Providers
Deepgram
Deepgram is a speech-to-text API. OpenClaw uses it for inbound audio/voice-note
transcription through tools.media.audio and for Voice Call streaming STT
through plugins.entries.voice-call.config.streaming.
Batch transcription uploads the complete audio file to Deepgram. Flux models
use Deepgram's one-shot WebSocket instead. Both paths inject the transcript
into the reply pipeline ({{Transcript}} + [Audio] block).
Voice Call streaming forwards live G.711 u-law frames over Deepgram's
WebSocket listen endpoint and emits partial/final transcripts as Deepgram
returns them.
| Detail | Value |
|---|---|
| Docs | developers.deepgram.com |
| Auth | DEEPGRAM_API_KEY |
| Default model | nova-3 |
Getting started
Set your API key
DEEPGRAM_API_KEY=dg_...Enable the audio provider
{ tools: { media: { models: [{ provider: "deepgram", model: "nova-3", capabilities: ["audio"] }], audio: { enabled: true, }, }, },}Send a voice note
Send an audio message through any connected channel. OpenClaw transcribes it via Deepgram and injects the transcript into the reply pipeline.
Configuration options
| Option | Path | Description |
|---|---|---|
model |
tools.media.models[].model |
Deepgram model id (default: nova-3) |
language |
tools.media.models[].language |
Language hint (optional) |
providerOptions.deepgram merges extra query params directly into the
Deepgram /listen request, so any Deepgram-supported param name works
(for example detect_language, punctuate, smart_format):
With language hint
{ tools: { media: { models: [ { provider: "deepgram", model: "nova-3", language: "en", capabilities: ["audio"] }, ], audio: { enabled: true, }, }, },}With Deepgram options
{ tools: { media: { models: [{ provider: "deepgram", model: "nova-3", capabilities: ["audio"] }], audio: { enabled: true, providerOptions: { deepgram: { detect_language: true, punctuate: true, smart_format: true, }, }, }, }, },}Flux models
Use flux-general-en or flux-general-multi for Deepgram Flux. OpenClaw
converts the voice note to 16 kHz mono linear16 audio with ffmpeg, then sends
it to Deepgram's /v2/listen WebSocket endpoint.
Set tools.media.models[].model to either Flux model in the getting-started
configuration above.
Flux supports eager_eot_threshold, eot_threshold, eot_timeout_ms,
keyterm, language_hint, mip_opt_out, numerals, profanity_filter,
redact, and tag in providerOptions.deepgram. OpenClaw ignores batch-only
options such as detect_language, punctuate, and smart_format on Flux.
Language hints apply only to flux-general-multi: OpenClaw maps the model entry's
language setting to language_hint, with an explicit provider option taking
precedence. Both settings are ignored for the English-only flux-general-en model.
Voice Call streaming STT
The bundled deepgram plugin also registers a realtime transcription provider
for the Voice Call plugin.
| Setting | Config path | Default |
|---|---|---|
| API key | plugins.entries.voice-call.config.streaming.providers.deepgram.apiKey |
Falls back to DEEPGRAM_API_KEY |
| Base URL | ...deepgram.baseUrl |
DEEPGRAM_BASE_URL or Deepgram's public API |
| Model | ...deepgram.model |
nova-3 |
| Language | ...deepgram.language |
(unset) |
| Encoding | ...deepgram.encoding |
mulaw |
| Sample rate | ...deepgram.sampleRate |
8000 |
| Endpointing | ...deepgram.endpointingMs |
800 |
| Interim results | ...deepgram.interimResults |
true |
{ plugins: { entries: { "voice-call": { config: { streaming: { enabled: true, provider: "deepgram", providers: { deepgram: { apiKey: "${DEEPGRAM_API_KEY}", model: "nova-3", endpointingMs: 800, language: "en-US", }, }, }, }, }, }, },}For a Deepgram custom endpoint,
set baseUrl to the endpoint root, including any base path but not /listen.
Realtime endpoints accept http://, https://, ws://, and wss://. HTTP
maps to WS, HTTPS maps to WSS, and explicit WebSocket schemes stay unchanged.
Malformed URLs and other schemes fail during session setup.
Notes
Authentication
Authentication follows the standard provider auth order. DEEPGRAM_API_KEY is
the simplest path.
Proxy and custom endpoints
Override endpoints or headers on the Deepgram tools.media.models[] entry when using a proxy.
Output behavior
Output follows the same audio rules as other providers (size caps, timeouts, transcript injection).
Flux requires ffmpeg
Install ffmpeg with the gateway host's package manager before selecting a
Flux model.