Files
Hermes-Skills/devops/hermes-voice-setup/SKILL.md
T
Hermes Skills Manager 6770bc9b9d Initial commit: Hermes Agent skills collection
- Trading skills (OKX, dividend, lottery, quantitative)
- Creative skills (ASCII art, diagrams, video)
- Development skills (GitHub, debugging, TDD)
- Research skills (arXiv, blog monitoring)
- Productivity skills (email, documents, notes)
- MCP integration skills
- Custom user skills
2026-07-05 02:31:15 -04:00

4.4 KiB

name, description, tags
name description tags
hermes-voice-setup Configure Hermes Agent voice mode — STT (speech-to-text) and TTS (text-to-speech) for CLI, Telegram, and Discord. Covers faster-whisper, Edge TTS, Groq Whisper, ElevenLabs, and Chinese/multilingual voice selection.
hermes
voice
stt
tts
whisper
telegram
speech

Hermes Voice Mode Setup

Configure Hermes Agent for voice interaction: speech-to-text (STT) input and text-to-speech (TTS) output across CLI, Telegram, and Discord.

Quick Reference

Component Free Option Premium Option
STT faster-whisper (local) Groq Whisper / OpenAI Whisper
TTS Edge TTS (free, no key) ElevenLabs / OpenAI TTS

Step 1: Install Dependencies

# System deps (Ubuntu/Debian)
sudo apt install portaudio19-dev ffmpeg libopus0 espeak-ng

# Python deps — in Hermes venv
source ~/.hermes/hermes-agent/venv/bin/activate
pip install faster-whisper
pip install edge-tts

Mac:

brew install portaudio ffmpeg opus espeak-ng

Step 2: Configure STT

# Local Whisper (free, no API key — recommended)
hermes config set stt.enabled true
hermes config set stt.provider local
hermes config set stt.local.model base    # tiny/base/small/medium/large-v3

Model sizes for faster-whisper:

Model Size Speed Accuracy Best For
tiny ~75MB Fastest Low Quick tests
base ~150MB Fast Good Default, balances speed/quality
small ~500MB Medium Better Multilingual
medium ~1.5GB Slow High Production Chinese/English
large-v3 ~3GB Slowest Best Maximum accuracy

For Chinese voice input, medium or large-v3 significantly outperform base.

Step 3: Configure TTS

# Edge TTS (free, no API key)
hermes config set tts.provider edge

Chinese Voice Options (Edge TTS)

Voice Gender Style Best For
zh-CN-XiaoxiaoNeural Female Warm Default Chinese voice
zh-CN-XiaoyiNeural Female Lively Casual/fun
zh-CN-YunxiNeural Male Sunshine General male voice
zh-CN-YunyangNeural Male Professional News/formal
zh-CN-YunjianNeural Male Passionate Sports/energetic
zh-CN-YunxiaNeural Male Cute Light content
hermes config set tts.edge.voice "zh-CN-XiaoxiaoNeural"

List Available Voices

edge-tts --list-voices | grep zh-CN    # Chinese
edge-tts --list-voices | grep en-US    # English
edge-tts --list-voices                 # All languages

Step 4: Usage

CLI

/voice on       # Enable voice mode
/voice off      # Disable
/voice tts      # Toggle TTS output (hear spoken replies)
/voice status   # Show current state
# Ctrl+B to start/stop recording

Telegram

Voice messages sent to the bot are auto-transcribed. No special config needed — STT handles it.

To make the bot reply with voice:

/voice tts      # In Telegram chat

Discord

  • Voice messages in text channels: auto-transcribed
  • Voice channel: /voice join to have bot participate in VC

Pitfalls

  • faster-whisper model downloads on first use: The model (~150MB for base) downloads from HuggingFace on first transcription. If HuggingFace is blocked (China), pre-download or use a mirror. First transcription will be slow.
  • PortAudio not needed for Telegram: portaudio19-dev is only required for CLI microphone input (Ctrl+B). Telegram voice messages work without it since they receive pre-recorded audio.
  • Edge TTS requires internet: Edge TTS calls Microsoft's cloud API. It's free but needs network access. For fully offline TTS, use NeuTTS: pip install neutts[all].
  • Chinese STT accuracy: base model works for English but struggles with Chinese. Upgrade to medium or large-v3 for better Chinese recognition: hermes config set stt.local.model medium.
  • TTS voice change needs gateway restart: After changing tts.edge.voice, restart the gateway for Telegram/Discord: hermes gateway restart. CLI picks it up on next session.
  • edge-tts vs edge_tts: The pip package is edge-tts (with hyphen). The Python import is edge_tts (with underscore). Don't confuse them.
  • Groq Whisper for fast cloud STT: If local Whisper is too slow, set GROQ_API_KEY in ~/.hermes/.env and use hermes config set stt.provider groq. Free tier has generous limits.