Files
Hermes-Skills/devops/hermes-voice-setup/SKILL.md
T
Hermes Skills Manager 6770bc9b9d Initial commit: Hermes Agent skills collection
- Trading skills (OKX, dividend, lottery, quantitative)
- Creative skills (ASCII art, diagrams, video)
- Development skills (GitHub, debugging, TDD)
- Research skills (arXiv, blog monitoring)
- Productivity skills (email, documents, notes)
- MCP integration skills
- Custom user skills
2026-07-05 02:31:15 -04:00

117 lines
4.4 KiB
Markdown

---
name: hermes-voice-setup
description: "Configure Hermes Agent voice mode — STT (speech-to-text) and TTS (text-to-speech) for CLI, Telegram, and Discord. Covers faster-whisper, Edge TTS, Groq Whisper, ElevenLabs, and Chinese/multilingual voice selection."
tags: [hermes, voice, stt, tts, whisper, telegram, speech]
---
# Hermes Voice Mode Setup
Configure Hermes Agent for voice interaction: speech-to-text (STT) input and text-to-speech (TTS) output across CLI, Telegram, and Discord.
## Quick Reference
| Component | Free Option | Premium Option |
|-----------|-------------|----------------|
| **STT** | faster-whisper (local) | Groq Whisper / OpenAI Whisper |
| **TTS** | Edge TTS (free, no key) | ElevenLabs / OpenAI TTS |
## Step 1: Install Dependencies
```bash
# System deps (Ubuntu/Debian)
sudo apt install portaudio19-dev ffmpeg libopus0 espeak-ng
# Python deps — in Hermes venv
source ~/.hermes/hermes-agent/venv/bin/activate
pip install faster-whisper
pip install edge-tts
```
**Mac:**
```bash
brew install portaudio ffmpeg opus espeak-ng
```
## Step 2: Configure STT
```bash
# Local Whisper (free, no API key — recommended)
hermes config set stt.enabled true
hermes config set stt.provider local
hermes config set stt.local.model base # tiny/base/small/medium/large-v3
```
**Model sizes for faster-whisper:**
| Model | Size | Speed | Accuracy | Best For |
|-------|------|-------|----------|----------|
| tiny | ~75MB | Fastest | Low | Quick tests |
| base | ~150MB | Fast | Good | Default, balances speed/quality |
| small | ~500MB | Medium | Better | Multilingual |
| medium | ~1.5GB | Slow | High | Production Chinese/English |
| large-v3 | ~3GB | Slowest | Best | Maximum accuracy |
For Chinese voice input, `medium` or `large-v3` significantly outperform `base`.
## Step 3: Configure TTS
```bash
# Edge TTS (free, no API key)
hermes config set tts.provider edge
```
### Chinese Voice Options (Edge TTS)
| Voice | Gender | Style | Best For |
|-------|--------|-------|----------|
| `zh-CN-XiaoxiaoNeural` | Female | Warm | Default Chinese voice |
| `zh-CN-XiaoyiNeural` | Female | Lively | Casual/fun |
| `zh-CN-YunxiNeural` | Male | Sunshine | General male voice |
| `zh-CN-YunyangNeural` | Male | Professional | News/formal |
| `zh-CN-YunjianNeural` | Male | Passionate | Sports/energetic |
| `zh-CN-YunxiaNeural` | Male | Cute | Light content |
```bash
hermes config set tts.edge.voice "zh-CN-XiaoxiaoNeural"
```
### List Available Voices
```bash
edge-tts --list-voices | grep zh-CN # Chinese
edge-tts --list-voices | grep en-US # English
edge-tts --list-voices # All languages
```
## Step 4: Usage
### CLI
```
/voice on # Enable voice mode
/voice off # Disable
/voice tts # Toggle TTS output (hear spoken replies)
/voice status # Show current state
# Ctrl+B to start/stop recording
```
### Telegram
Voice messages sent to the bot are auto-transcribed. No special config needed — STT handles it.
To make the bot reply with voice:
```
/voice tts # In Telegram chat
```
### Discord
- Voice messages in text channels: auto-transcribed
- Voice channel: `/voice join` to have bot participate in VC
## Pitfalls
- **faster-whisper model downloads on first use**: The model (~150MB for `base`) downloads from HuggingFace on first transcription. If HuggingFace is blocked (China), pre-download or use a mirror. First transcription will be slow.
- **PortAudio not needed for Telegram**: `portaudio19-dev` is only required for CLI microphone input (Ctrl+B). Telegram voice messages work without it since they receive pre-recorded audio.
- **Edge TTS requires internet**: Edge TTS calls Microsoft's cloud API. It's free but needs network access. For fully offline TTS, use NeuTTS: `pip install neutts[all]`.
- **Chinese STT accuracy**: `base` model works for English but struggles with Chinese. Upgrade to `medium` or `large-v3` for better Chinese recognition: `hermes config set stt.local.model medium`.
- **TTS voice change needs gateway restart**: After changing `tts.edge.voice`, restart the gateway for Telegram/Discord: `hermes gateway restart`. CLI picks it up on next session.
- **edge-tts vs edge_tts**: The pip package is `edge-tts` (with hyphen). The Python import is `edge_tts` (with underscore). Don't confuse them.
- **Groq Whisper for fast cloud STT**: If local Whisper is too slow, set `GROQ_API_KEY` in `~/.hermes/.env` and use `hermes config set stt.provider groq`. Free tier has generous limits.