Skip to main content
Version: Latest

Start Session

Starts an interactive avatar session with a session token. Use livekit.url and livekit.token from the response to connect to real-time video streaming.

What is LiveKit?​

LiveKit is a WebRTC-based real-time video streaming stack. After Start Session returns livekit.url, livekit.token, and room_name, the client calls something like room.connect(url, token) to join the LiveKit server. Once connected, the avatar’s video, audio, and data channels stream in real time for a conversational experience.

POST /api/v2/sessions/start

Headers​

HeaderValue
AuthorizationBearer {session_token}
Content-Typeapplication/json

Body​

FieldTypeRequiredDescription
avatar_idstringrequiredAvatar / model ID
avatar_personaobjectoptionalPersona and LLM overrides. If omitted, server defaults apply: English (en), OpenAI, gpt-4.1-nano.
avatar_persona.languagestringoptionale.g. en, ko, en-US
avatar_persona.llm_configurationsobjectoptionalLLM provider/model settings. Server defaults apply if omitted.
avatar_persona.llm_configurations.providerstringoptionale.g. openai, anthropic, google, custom
avatar_persona.llm_configurations.modelstringoptionalModel ID
avatar_persona.llm_configurations.temperaturenumberoptionalTemperature forwarded to the LLM
avatar_persona.llm_configurations.custom_settingsobjectoptionalgreeting_text, system_prompt overrides
avatar_persona.llm_configurations.custom_httpobjectoptionalUsed only when provider is "custom". Requires endpoint or url.
max_session_durationintegeroptionalMaximum session length in minutes. Must be within your Plan limit.
lip_audio_modestringoptionalllm_tts or external_pcm. If omitted, the default audio behavior applies.

If max_session_duration is omitted, your Account subscription limit applies. max_session_numbers is not accepted in the request body; it is determined by your subscription plan.

Custom HTTP LLM​

Use avatar_persona.llm_configurations.provider: "custom" with custom_http when you want the session to use your own HTTP LLM endpoint.

FieldTypeDescription
endpointstringFull POST URL for your custom LLM server
urlstringAlias for endpoint; mapped to endpoint if endpoint is omitted
connect_timeout_sec, read_timeout_sec, timeout_secnumberConnection/read timeouts in seconds
max_messages, max_chars, max_buffer_charsnumberConversation trimming and text flush limits
http_error_messagestringMessage prefix yielded when the upstream returns HTTP 4xx/5xx
stream_connect_retries, stream_connect_retry_delay_secnumberRetry settings before the first SSE line or response body
headers, extra_headersobjectStatic headers. Prefer one; if both are present, headers is used.
auth_pluginsarrayAuth plugin objects
streambooleanDefaults to true for SSE. Use false for a single JSON response.
api_presetstringopenai_compatible (default), anthropic_messages, or gemini_generate_content

api_preset reference formats:

api_preset selects the HTTP payload shape and response parser used for your custom endpoint. It does not mean every option in the provider's official API is supported. If you need provider-specific features beyond the supported shape, put them behind your own HTTP gateway and return one of the supported response formats.

PresetBest forStreaming responseNon-streaming response
openai_compatibleOpenAI-compatible Chat Completions endpoints and most LLM gatewaysSSE chunks with text deltasOne JSON body with text in choices[0].message.content
anthropic_messagesAnthropic Messages-compatible endpointsSSE events with text deltasOne JSON body with text blocks in content
gemini_generate_contentGemini GenerateContent-compatible endpointsSSE response from streamGenerateContentOne JSON body from generateContent

For protocol details, supported customization, and socket/WebSocket limitations, see Custom HTTP LLM.

Behavior notes:

  • api_preset defaults to openai_compatible. Use anthropic_messages for Anthropic Messages API payloads and gemini_generate_content for Gemini GenerateContent payloads.
  • When api_preset is anthropic_messages, model is required.
  • stream defaults to true and expects SSE. Set stream: false only when the upstream returns one JSON response; for openai_compatible, text is read from choices[0].message.content.
  • Header merge order is default headers (Content-Type, Accept) → headers / extra_headers → auth_plugins; later values override earlier values.
  • Default timing values are connect_timeout_sec: 10, read_timeout_sec: 120, stream_connect_retries: 1, and stream_connect_retry_delay_sec: 0.4.
  • For production, prefer server-side secret assembly with auth_plugins or environment-backed headers instead of accepting raw API keys from untrusted clients.

Greeting / first utterance​

If custom_settings.greeting_text is present and non-empty after trim, the service automatically sends one server-side chat message immediately after Start Session succeeds.

  • Message: [FIRST_MESSAGE] + trimmed greeting_text
  • Server-sent Talk message type: first_message

Tip: The text appended after [FIRST_MESSAGE] is spoken by the avatar as-is. You can use this when you already have the exact text the avatar should say and want to inject that response directly.

When lip_audio_mode is external_pcm, the service does not send this automatic first message.

Examples​

Basic session​

curl https://ai-streamer.deepbrain.io/api/v2/sessions/start \
-X POST \
-H "Content-Type: application/json" \
-H "Authorization: Bearer ${SESSION_TOKEN}" \
-d '{
"avatar_id": "${YOUR_AVATAR_ID}"
}'

With custom LLM​

Connect a custom HTTP LLM endpoint.

curl https://ai-streamer.deepbrain.io/api/v2/sessions/start \
-X POST \
-H "Content-Type: application/json" \
-H "Authorization: Bearer ${SESSION_TOKEN}" \
-d '{
"avatar_id": "${YOUR_AVATAR_ID}",
"avatar_persona": {
"language": "en",
"llm_configurations": {
"provider": "custom",
"model": "my-llm-model",
"custom_settings": {
"system_prompt": "You are a helpful assistant.",
"greeting_text": "Hello! How can I help you today?"
},
"custom_http": {
"endpoint": "https://your-llm-backend.example.com/v1/chat/completions",
"api_preset": "openai_compatible",
"stream": true,
"read_timeout_sec": 180,
"headers": {
"Authorization": "Bearer ${YOUR_LLM_API_KEY}"
}
}
}
}
}'

Success (201, code: 1000)​

{
"code": 1000,
"data": {
"session_id": "${SESSION_ID}",
"max_session_duration": 20,
"max_session_numbers": 10,
"use_server_ready": true,
"livekit": {
"url": "wss://your-livekit-host",
"room_name": "room-…",
"token": "${LIVEKIT_ACCESS_TOKEN}"
}
},
"message": "Session created successfully"
}

When the request includes lip_audio_mode: "external_pcm", the response data also includes lip_audio for the LiveKit byte stream topic, PCM format, and segment behavior. If lip_audio_mode is omitted or llm_tts, lip_audio is omitted.

use_server_ready is a client coordination flag for LiveKit server-side readiness; it is currently returned as true for V2 starts.

Next step​

After Start Session, connect to LiveKit with livekit.url and livekit.token to stream the avatar’s video and audio in real time. If the session stays idle without user messages, call the Session Keep-Alive API periodically.