> ## Documentation Index
> Fetch the complete documentation index at: https://wholly.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Speech Recognition

> Transcribe audio to text using Whisper and other speech recognition models.

WhollyAPI hosts [Whisper](https://github.com/openai/whisper) and other speech recognition models. Given an audio file, they produce transcribed text with per-sentence timestamps.

Browse [all speech recognition models](https://whollyapi.com/models?category=speech-recognition).

## Models

* `openai/whisper-large` — best accuracy
* `openai/whisper-medium`, `openai/whisper-small`, `openai/whisper-base` — faster, lighter
* `openai/whisper-timestamped-medium` — per-word timestamp segmentation

## Example

<CodeGroup>
  ```bash cURL theme={null}
  curl -X POST \
      -H "Authorization: Bearer $WHOLLYAPI_TOKEN" \
      -F audio=@audio.mp3 \
      'https://whollyapi.com/api/v1/inference/openai/whisper-large'
  ```
</CodeGroup>

## Supported audio formats

* `mp3`
* `wav`

## Response

```json theme={null}
{
  "text": "Hello, this is a transcription of the audio file.",
  "segments": [
    {
      "start": 0.0,
      "end": 3.5,
      "text": "Hello, this is a transcription of the audio file."
    }
  ]
}
```

## Additional parameters

Each model exposes different parameters (language, task, etc.). Check the model's API documentation page for details.
