The AI Providers section lets you create one or more configurations for each AI engine, specifying the model to use, whether transcriptions are enabled, the response temperature, and other engine-specific settings. You can create as many configurations as you need; each one is later selectable as the provider for your AI Conversational Agents or for AI AMD.

SIP Caller currently supports the following engines:
To access AI Providers, go to "Settings" > "AI Providers".

Here you can configure the following fields:
- API Key: authenticates your requests to the Azure AI service. Refer to the Azure AI setup guide for instructions on how to obtain it.
- Endpoint: the URL of your Azure AI resource. Refer to the Azure AI setup guide for instructions on how to obtain it.
- Model: the Azure Voice Live model to use for AI agents. You can choose between Voice Live Pro (highest quality, best for complex conversational experiences), Voice Live Basic (balanced quality and cost), or Voice Live Lite (most cost-effective for simpler interactions). Refer to Azure's Voice Live pricing page for a full comparison of each tier.
- Enable Transcriptions: enable or disable call transcriptions for AI agents using this provider.
- Transcription Model: the model used to transcribe calls, available when transcriptions are enabled.
- Response Temperature: controls the creativity and variability of the AI's responses. Lower values produce more consistent, predictable replies; higher values allow for more varied and creative responses.
- Playback Speed: controls how fast the AI agent speaks. Adjust this to make the voice sound more natural for your use case.

Here you can configure the following fields:
- API Key: authenticates your requests to the Gemini AI service. Refer to the Gemini AI setup guide for instructions on how to obtain it.
- Model: the Gemini model to use for AI agents.
- Enable Transcriptions: enable or disable call transcriptions for AI agents using this provider.
- Response Temperature: controls the creativity and variability of the AI's responses. Lower values produce more consistent, predictable replies; higher values allow for more varied and creative responses.