# TensorX Speech API

Live and file speech-to-text on open models, hosted on NVIDIA B300 GPUs in the EU. This is a proof-of-concept environment.

- Base URL: `https://speech.tensorx.ai`
- Human-readable docs: `https://speech.tensorx.ai/docs`
- This file: `https://speech.tensorx.ai/docs.md` (the complete reference; nothing is left out of it)

## Authentication

Create a key on the **API keys** page of the app. Keys start with `txs_` and are shown once.

Send the key in one of these ways (first match wins):

| Where | Example |
|---|---|
| `Authorization` header | `Authorization: Bearer txs_...` |
| Azure OpenAI style header | `api-key: txs_...` |
| Query string (WebSocket clients that cannot set headers) | `?api_key=txs_...` |

A missing, revoked or wrong key gives HTTP 401, or on a WebSocket an error event followed by close code 1008.
A valid key whose owner does not have the client API switched on gives HTTP 403 with
`{"detail": "...", "code": "feature_disabled"}`, or the same error event and close code 1008 on a WebSocket.

## Data handling

Zero data retention. Audio and transcripts pass through memory only and are never written to disk or a database. For usage we store call ID, timings, audio duration, model, status and error code; never audio or text.

## Limits (POC)

| Limit | Value |
|---|---|
| Calls at once per key | 20 (stereo counts as 2). Ask us to raise it |
| Call length | 90 minutes |
| File size | 100 MB |
| Lines per model | Shared across all customers; `GET /v1/models` shows `in_use` and `max_calls` |

## Models

`GET /v1/models` (auth required) returns:

```json
{"object": "list", "data": [
  {"id": "nemotron-3.5-asr-streaming", "object": "model", "owned_by": "tensorx",
   "label": "Nemotron 3.5 ASR Streaming 0.6B", "languages": "40 languages and locales",
   "latency": "Text about 0.3 s behind speech", "max_calls": 400, "in_use": 3, "status": "ready"}
]}
```

`status` is `ready` or `offline`.

| Model ID | Languages | Text behind speech | Notes |
|---|---|---|---|
| `nemotron-3.5-asr-streaming` | 40 languages and locales | about 0.3 s | Default model |
| `nemotron-speech-streaming-en` | English | about 0.3 s | |
| `parakeet-unified-en` | English | about 1.1 s | |
| `voxtral-mini-4b-realtime` | 13 languages | about 0.8 s | |
| `kyutai-stt-2.6b-en` | English | about 2.5 s | Waits on purpose for accuracy |
| `kyutai-stt-1b-en-fr` | English, French | about 1 s | |

If `model` is omitted anywhere, `nemotron-3.5-asr-streaming` is used.

## Live option 1: OpenAI Realtime-compatible WebSocket

For clients already using OpenAI or Azure OpenAI realtime transcription: change the URL and key, keep the events.

```
wss://speech.tensorx.ai/v1/realtime?model=nemotron-3.5-asr-streaming
```

The model comes from, in order: the `model` query parameter, the model in the session update, then the default. OpenAI model names (for example `gpt-4o-transcribe`) map to the default model.

### Client events

| Event | Effect |
|---|---|
| `transcription_session.update` (beta) or `session.update` (GA) | Sets the input format. Must come before audio if you need a format other than `pcm16` 24 kHz. Answered with `transcription_session.updated` / `session.updated` |
| `input_audio_buffer.append` | `{"type":"input_audio_buffer.append","audio":"<base64>"}`. Send whole samples |
| `input_audio_buffer.commit` | Optional. Ends the current turn: everything heard so far is completed as one item, and the next audio starts a new item |
| `input_audio_buffer.clear` | Drops the current turn's audio without a transcript. Answered with `input_audio_buffer.cleared` |
| `session.close` | Ends the session. Closing the socket also ends it |

Any other event type gets an `error` with code `unsupported_event`; the session continues.

### Input formats

| Beta `input_audio_format` | GA `audio.input.format.type` | Audio |
|---|---|---|
| `pcm16` (default) | `audio/pcm` (set `rate`, default 24000) | 16-bit little-endian PCM, mono |
| `g711_ulaw` | `audio/pcmu` | G.711 μ-law, 8 kHz |
| `g711_alaw` | `audio/pcma` | G.711 A-law, 8 kHz |

Beta shape: `{"type":"transcription_session.update","session":{"input_audio_format":"pcm16"}}`
GA shape: `{"type":"session.update","session":{"audio":{"input":{"format":{"type":"audio/pcm","rate":16000}}}}}`

### Server events

| Event | When |
|---|---|
| `transcription_session.created` | On connect |
| `conversation.item.input_audio_transcription.delta` | New words while people talk: `item_id`, `content_index`, `delta` |
| `input_audio_buffer.committed` | A segment ended (sentence end found by the model, or your commit): `item_id`, `previous_item_id` |
| `conversation.item.input_audio_transcription.completed` | Final text of that item: `item_id`, `content_index`, `transcript` |
| `error` | `{"type":"error","error":{"type":"invalid_request_error","code":"...","message":"..."}}` |

The model finds sentence ends itself, so you get completed items without sending `commit`.

### Example (Python)

```python
import asyncio, base64, json, websockets

async def main():
    url = "wss://speech.tensorx.ai/v1/realtime?model=nemotron-3.5-asr-streaming"
    async with websockets.connect(url, additional_headers={"Authorization": "Bearer txs_..."}) as ws:
        await ws.send(json.dumps({"type": "transcription_session.update",
                                  "session": {"input_audio_format": "pcm16"}}))
        # send 24 kHz mono PCM16 chunks of about 100 ms:
        # await ws.send(json.dumps({"type": "input_audio_buffer.append",
        #                           "audio": base64.b64encode(chunk).decode()}))
        async for msg in ws:
            ev = json.loads(msg)
            if ev["type"].endswith("transcription.completed"):
                print(ev["transcript"])

asyncio.run(main())
```

## Live option 2: TensorX stream WebSocket (simplest)

```
wss://speech.tensorx.ai/v1/stream?model=nemotron-3.5-asr-streaming&sample_rate=16000&encoding=pcm_s16le&channels=1
```

| Query parameter | Values | Default |
|---|---|---|
| `model` | a model ID | `nemotron-3.5-asr-streaming` |
| `encoding` | `pcm_s16le`, `mulaw`, `alaw` | `pcm_s16le` |
| `sample_rate` | 8000 to 48000 | 16000 |
| `channels` | 1, or 2 for agent and caller | 1 |

- Send raw audio as binary frames, about 100 ms each. With `channels=2`, send interleaved stereo; every event carries `channel` 0 or 1.
- Send the text frame `{"type":"stop"}` to finish. You get the last text, then `done`, then the socket closes.

Server events (JSON text frames):

```json
{"type": "ready", "call_id": "call_...", "model": "nemotron-3.5-asr-streaming", "channels": 1, "sample_rate": 16000}
{"type": "partial", "channel": 0, "text": "I can see there's an outstanding", "audio_seconds": 3.2}
{"type": "final", "channel": 0, "text": "I can see there's an outstanding balance on your account.", "audio_seconds": 4.1}
{"type": "done", "call_id": "call_...", "usage": {"audio_seconds": 70.2, "channels": 1, "first_text_ms": 640, "final_ms": 310}}
{"type": "error", "code": "model_busy", "message": "..."}
```

`partial` replaces the previous partial for that channel; `final` is settled text. Close codes: 1000 normal, 1008 bad request or key, 1011 call failed, 1013 busy or over a limit (retry later).

## Files: OpenAI-compatible upload

```
POST https://speech.tensorx.ai/v1/audio/transcriptions
Content-Type: multipart/form-data
```

| Field | Required | Notes |
|---|---|---|
| `file` | yes | wav, mp3, m4a, ogg/opus, flac and most other formats |
| `model` | no | a model ID |
| `response_format` | no | `json` (default), `text`, `verbose_json` |

```bash
curl https://speech.tensorx.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer txs_..." \
  -F model=nemotron-3.5-asr-streaming \
  -F file=@call.wav \
  -F response_format=json
```

Responses:

- `json`: `{"text": "..."}`
- `text`: the transcript as plain text
- `verbose_json`: `{"task": "transcribe", "duration": 71.8, "text": "...", "model": "...", "segments": [{"id": 0, "text": "..."}], "processing_seconds": 5.4}`

## Errors

| Code | HTTP status (files) | Meaning |
|---|---|---|
| `unauthorised` / `invalid_api_key` | 401 | Missing, revoked or wrong key |
| `feature_disabled` | 403 | The key is valid, but the client API is switched off for its owner's account. Ask your TensorX contact |
| `access_check_unavailable` | 503 | Access could not be checked just now, so the call was refused. Retry with backoff (WebSocket close code 1013) |
| `unknown_model` | 503 | The model ID is not in `/v1/models` |
| `model_busy` | 429 | All lines for that model are in use. Retry with backoff |
| `key_limit` | 429 | Your key is at its calls-at-once limit |
| `model_unavailable` | 503 | The model server is not reachable right now |
| `model_error` | 502 | The model stopped during the call or did not finish the transcript |
| `call_too_long` | 503 | The call passed the POC length limit |
| `bad_request` / `invalid_request_error` | 400 | Bad parameters, unreadable audio or unsupported format |
| `unsupported_event` | n/a | Realtime only: an event type this endpoint does not handle |
| `internal_error` / `server_error` | n/a | Unexpected failure; the call is ended |

Files over the size limit get HTTP 413.
