- Trading skills (OKX, dividend, lottery, quantitative) - Creative skills (ASCII art, diagrams, video) - Development skills (GitHub, debugging, TDD) - Research skills (arXiv, blog monitoring) - Productivity skills (email, documents, notes) - MCP integration skills - Custom user skills
4.4 KiB
4.4 KiB
name, description, tags
| name | description | tags | |||||||
|---|---|---|---|---|---|---|---|---|---|
| hermes-voice-setup | Configure Hermes Agent voice mode — STT (speech-to-text) and TTS (text-to-speech) for CLI, Telegram, and Discord. Covers faster-whisper, Edge TTS, Groq Whisper, ElevenLabs, and Chinese/multilingual voice selection. |
|
Hermes Voice Mode Setup
Configure Hermes Agent for voice interaction: speech-to-text (STT) input and text-to-speech (TTS) output across CLI, Telegram, and Discord.
Quick Reference
| Component | Free Option | Premium Option |
|---|---|---|
| STT | faster-whisper (local) | Groq Whisper / OpenAI Whisper |
| TTS | Edge TTS (free, no key) | ElevenLabs / OpenAI TTS |
Step 1: Install Dependencies
# System deps (Ubuntu/Debian)
sudo apt install portaudio19-dev ffmpeg libopus0 espeak-ng
# Python deps — in Hermes venv
source ~/.hermes/hermes-agent/venv/bin/activate
pip install faster-whisper
pip install edge-tts
Mac:
brew install portaudio ffmpeg opus espeak-ng
Step 2: Configure STT
# Local Whisper (free, no API key — recommended)
hermes config set stt.enabled true
hermes config set stt.provider local
hermes config set stt.local.model base # tiny/base/small/medium/large-v3
Model sizes for faster-whisper:
| Model | Size | Speed | Accuracy | Best For |
|---|---|---|---|---|
| tiny | ~75MB | Fastest | Low | Quick tests |
| base | ~150MB | Fast | Good | Default, balances speed/quality |
| small | ~500MB | Medium | Better | Multilingual |
| medium | ~1.5GB | Slow | High | Production Chinese/English |
| large-v3 | ~3GB | Slowest | Best | Maximum accuracy |
For Chinese voice input, medium or large-v3 significantly outperform base.
Step 3: Configure TTS
# Edge TTS (free, no API key)
hermes config set tts.provider edge
Chinese Voice Options (Edge TTS)
| Voice | Gender | Style | Best For |
|---|---|---|---|
zh-CN-XiaoxiaoNeural |
Female | Warm | Default Chinese voice |
zh-CN-XiaoyiNeural |
Female | Lively | Casual/fun |
zh-CN-YunxiNeural |
Male | Sunshine | General male voice |
zh-CN-YunyangNeural |
Male | Professional | News/formal |
zh-CN-YunjianNeural |
Male | Passionate | Sports/energetic |
zh-CN-YunxiaNeural |
Male | Cute | Light content |
hermes config set tts.edge.voice "zh-CN-XiaoxiaoNeural"
List Available Voices
edge-tts --list-voices | grep zh-CN # Chinese
edge-tts --list-voices | grep en-US # English
edge-tts --list-voices # All languages
Step 4: Usage
CLI
/voice on # Enable voice mode
/voice off # Disable
/voice tts # Toggle TTS output (hear spoken replies)
/voice status # Show current state
# Ctrl+B to start/stop recording
Telegram
Voice messages sent to the bot are auto-transcribed. No special config needed — STT handles it.
To make the bot reply with voice:
/voice tts # In Telegram chat
Discord
- Voice messages in text channels: auto-transcribed
- Voice channel:
/voice jointo have bot participate in VC
Pitfalls
- faster-whisper model downloads on first use: The model (~150MB for
base) downloads from HuggingFace on first transcription. If HuggingFace is blocked (China), pre-download or use a mirror. First transcription will be slow. - PortAudio not needed for Telegram:
portaudio19-devis only required for CLI microphone input (Ctrl+B). Telegram voice messages work without it since they receive pre-recorded audio. - Edge TTS requires internet: Edge TTS calls Microsoft's cloud API. It's free but needs network access. For fully offline TTS, use NeuTTS:
pip install neutts[all]. - Chinese STT accuracy:
basemodel works for English but struggles with Chinese. Upgrade tomediumorlarge-v3for better Chinese recognition:hermes config set stt.local.model medium. - TTS voice change needs gateway restart: After changing
tts.edge.voice, restart the gateway for Telegram/Discord:hermes gateway restart. CLI picks it up on next session. - edge-tts vs edge_tts: The pip package is
edge-tts(with hyphen). The Python import isedge_tts(with underscore). Don't confuse them. - Groq Whisper for fast cloud STT: If local Whisper is too slow, set
GROQ_API_KEYin~/.hermes/.envand usehermes config set stt.provider groq. Free tier has generous limits.