Skip to content

Latest commit

 

History

History
 
 

Folders and files

README.md

Overview

A pip-installable Python SDK for connecting to Omi wearable devices over Bluetooth, decoding Opus-encoded audio, and transcribing it in real time using Deepgram.

Deepgram transcription requires websockets 14.0 or newer. The SDK's dependency declarations enforce this minimum for its authenticated WebSocket handshake.

Connect to any Omi device Decode Opus audio to PCM Deepgram-powered STT

Quick Start

Get a free API key from [Deepgram](https://deepgram.com):
```bash
export DEEPGRAM_API_KEY=your_actual_deepgram_key
```
Scan for nearby Bluetooth devices:
```bash
omi-scan
```

Look for a device named "Omi" and copy its MAC address:

```
0. Omi [7F52EC55-50C9-D1B9-E8D7-19B83217C97D]
```
```python import asyncio import os from omi import listen_to_omi, OmiOpusDecoder, transcribe from asyncio import Queue
# Configuration
OMI_MAC = "YOUR_OMI_MAC_ADDRESS_HERE"  # From omi-scan
OMI_CHAR_UUID = "19B10001-E8F2-537E-4F6C-D104768A1214"  # Standard Omi audio UUID
DEEPGRAM_API_KEY = os.getenv("DEEPGRAM_API_KEY")

async def main():
    audio_queue = Queue()
    decoder = OmiOpusDecoder()

    def handle_audio(sender, data):
        pcm_data = decoder.decode_packet(data)
        if pcm_data:
            audio_queue.put_nowait(pcm_data)

    def handle_transcript(transcript):
        print(f"Transcription: {transcript}")
        # Save to file, send to API, etc.

    # Start transcription and device connection
    await asyncio.gather(
        listen_to_omi(OMI_MAC, OMI_CHAR_UUID, handle_audio),
        transcribe(audio_queue, DEEPGRAM_API_KEY, on_transcript=handle_transcript)
    )

if __name__ == "__main__":
    asyncio.run(main())
```
```bash python examples/main.py ```
The example will:
- Connect to your Omi device via Bluetooth
- Decode incoming Opus audio packets to PCM
- Transcribe audio in real-time using Deepgram
- Print transcriptions to the console

Deepgram language selection

The Deepgram engine defaults to en-US. Set language to a language supported by your chosen Deepgram model when transcribing other languages:

from omi.stt import create_transcriber

transcriber = create_transcriber("deepgram", api_key=DEEPGRAM_API_KEY, language="es")
await transcriber.run(audio_queue, on_transcript=handle_transcript)

The omi.transcribe.transcribe function also forwards language="es" through its engine options. This selects a transcription language; it does not translate audio or enable automatic language detection. See Deepgram's model and language support.

Development

Stopping Deepgram transcription

Cancel and await the task running the transcriber to stop it. The SDK stops sending PCM, sends Deepgram's CloseStream, and keeps receiving final transcripts until the server closes or the receive timeout expires. drain_timeout bounds both the terminal send and the receive wait (default: five seconds per phase). Callbacks can therefore run while cancellation is being awaited. Set drain_timeout=0 on create_transcriber("deepgram", ...) to skip waiting for final results; the terminal send still has a five-second timeout. A server EOF or connection error follows the reconnect path; it does not send CloseStream.

Local Development Setup

git clone https://github.com/BasedHardware/omi.git
cd omi/sdks/python

# Create virtual environment
python3 -m venv venv
source venv/bin/activate

# Install in editable mode
pip install -e .

# Install dev dependencies
pip install -e ".[dev]"

Run the SDK's hardware-free regression suite with bash test.sh. It uses the repository's uv environment manager to test the minimum supported websockets version and is also selected by the shared local/CI check manifest.

License

MIT License — built by the community.