Skip to content

Supported models

The complete list of providers and models available today, grouped by what they're used for. Defaults are marked.

Current at time of writing

Models change often. This page is a snapshot — the live model registry in Kiwi is always the source of truth.

Chat & reasoning

Provider Models
OpenAI gpt-4o, gpt-4o-mini, gpt-5, gpt-5-mini (default), gpt-5-nano, gpt-5.1, gpt-5.1-codex, gpt-5.2, gpt-5.2-codex, gpt-5.3-codex, gpt-5.4, gpt-5.4-nano, o3, o3-mini, o4-mini
Google Gemini gemini-2.5-flash, gemini-2.5-pro, gemini-3-flash-preview, gemini-3.1-pro-preview

Reasoning depth is configurable on models that support it.

Image generation

Provider Models
Google nano-banana-2 (default), nano-banana-2-lite, nano-banana-pro
OpenAI gpt-image-1.5, gpt-image-2
Adobe Firefly firefly-5

Options include aspect ratio, size (1K/2K/4K, Google models) and quality (medium/high/auto, OpenAI models). Omitted settings use the effective defaults below.

Image editing

Provider Models
Google nano-banana-2 (default), nano-banana-2-lite, nano-banana-pro
OpenAI gpt-image-1.5
Adobe Firefly firefly-5

Video generation

Provider Models
Google Veo / Gemini Video gemini-omni-1.1-flash (default), veo-3.1-fast-generate-001, veo-3.1-generate-001
Runway (Seedance variants) seedance2, seedance2_fast, seedance2_mini

Options include aspect ratio (9:16, 16:9), duration (3–15s), resolution (360p/480p/720p/1080p/4K) and frame animation. Exact choices depend on the selected model.

Audio generation

Type Provider Models
Music Google Lyria lyria-002, lyria-3-clip-preview (default), lyria-3-pro-preview
Sound effects ElevenLabs default sound-effect model
Speech ElevenLabs eleven_v3 (default), eleven_multilingual_v2, eleven_flash_v2_5, eleven_turbo_v2_5

Effective media defaults

When a setting is omitted, Kiwi applies these defaults for the selected model. Explicit requested settings take priority when the model supports them.

Operation Model family Effective defaults
Image generation Google Nano Banana 1:1, 1K.
Image generation OpenAI 1:1, quality=auto.
Image generation Firefly 5 1:1 for text-only generation.
Video generation Google Veo 16:9, 8 seconds, 1080p.
Video generation Seedance 16:9, 5 seconds, 480p.
Video generation Omni 16:9, 5 seconds, 720p for an initial generation.
Sound effects ElevenLabs 5 seconds.
Speech ElevenLabs Roger voice (CwhRBWXzGAHq8TQ4Fs17).

Conditioning can replace a shape default. Google and OpenAI image generation can derive shape from reference images when no ratio was requested; Firefly edits use their references; Omni follow-ups use their previous interaction; Visual Edit preserves the source image's framing.

Captioning

Capability Providers
Image & video captioning OpenAI, Google Gemini (default)

Provider notes

  • Anthropic (Claude) — coming soon as a new chat provider.
  • Azure OpenAI — not planned.
  • gpt-image-2 is available for generation only, not editing.

Related: Model · Generate media · Coming soon