Skip to content

Model capability, tiers and roadmap ​

As of 2026-08-12. Written to answer a partner's three questions — what the platform can do, what is coming when, and whether Auto routing is ours to configure — and to plan the video expansion behind it.

Read the honest summary first, because the headline number is easy to misquote.

The one number that matters ​

CountWhat it means
Model rows in the catalogue72Integrated in code: adapter written, request/response mapped
Servable right now29Model enabled and its provider has a working API key
Verified by a real call13We have actually generated something with it and logged the result
Verified failing6Real call attempted and refused — needs a fix or retirement
Unverified43Never called. Assume nothing about these until they are verified

Do not tell a partner "72 models". Say 29 live, 13 independently verified, 72 integrated — the gap is not engineering, it is API keys. Only three providers are keyed today (Google/Gemini, OpenAI, ElevenLabs). The other fifteen have finished adapters sitting behind a missing credential, which is why the platform reports them as unservable rather than pretending otherwise.

That is the whole story of this roadmap: most of what a partner is asking for is a procurement task measured in days, not a build measured in months.

1. Feature-wise capability, by tier ​

Tier is quality-and-cost banding, not vendor prestige:

  • Flagship — frontier output, highest unit cost. What you demo.
  • Premium — strong, production-grade, mid cost. What most paid usage should land on.
  • Standard — fast and cheap. What free and entry tiers run on.

LIVE = servable today · KEY = integrated, waiting on a provider key.

Text generation & chat ​

TierModelStatus
FlagshipClaude Opus 4.7 · Gemini 3 ProKEY · LIVE (Gemini 3 Pro currently failing verification)
PremiumGemini 3.5 Flash · Gemini Pro (latest) · GPT-4o · Claude Sonnet 4.6 · Kimi K2.6LIVE ×3, KEY ×2
StandardGemini Flash Lite · GPT-4o mini · Gemma 4 31B · DeepSeek Chat · Grok 2LIVE ×3, KEY ×2

Vision / multimodal understanding ​

Same Gemini and GPT-4o line as above — 7 live. Image-in, text-out; used for describe/analyse/OCR-style features.

Image generation ​

TierModelStatus
FlagshipGPT Image 2 · Nano Banana Pro · FLUX 1.1 Pro · Ideogram v2LIVE ×2, KEY ×2
PremiumGPT Image 1.5 · Gemini 3.1 Flash Image · Gemini 2.5 Flash Image · Stable Diffusion 3 Large · Qwen ImageLIVE ×3, KEY ×2
StandardGPT Image 1 Mini · FLUX Schnell · FLUX Dev · SDXL · MiniMax Image 01LIVE ×1, KEY ×4

19 rows, 6 live. Also live: image-to-image, inpaint (mask), and an enhance/upscale path.

Video generation — the thin spot ​

TierModelStatus
FlagshipSora 2 Pro · Veo 3.1 · Runway Gen-4.5 · Kling 2.5 Turbo Pro · Seedance 2.0LIVE ×2, KEY ×3
PremiumSora 2 · Veo 3.1 Fast · Runway Gen-4 Turbo · Hailuo 2.3 · Wan 2.5 · Luma Ray 3.2 · Runway Aleph 2.0LIVE ×2, KEY ×5
StandardVeo 3.1 Lite · Kling 1.6 Pro · Runway Gen-3 · Luma Uni-1 · Hailuo (Runware) · Wan (Runware)LIVE ×1, KEY ×5

27 rows, 5 live (3 Veo + 2 Sora), none verified. Everything else is one API key away. Note Seedance is already integrated twice — as runware/seedance-via-runware and runway/seedance-2.0 — so the partner's ask for Seedance is a signup, not a sprint. (The second row is also mis-filed: Seedance is ByteDance's model, not Runway's. It needs correcting or deleting.)

Beyond text-to-video, already implemented: image-to-video, first-and-last-frame, video extend, and lip-sync (Kling), plus avatar video via Tavus. All KEY.

Speech, music and audio ​

TierModelStatus
FlagshipElevenLabs Multilingual v3 · Eleven MusicLIVE, both currently failing verification
PremiumMultilingual v2 · Turbo v2.5 · Gemini 3.1 Flash TTS · MiniMax Speech 2.8LIVE ×3, KEY ×1
StandardFlash v2.5 · Sound Effects · Gemini 2.5 Flash TTSLIVE ×3
Speech-to-textElevenLabs Scribe v1LIVE, failing verification

11 rows, 9 live — the strongest area, and the one with four verification failures to clear.

Also live (platform, not models) ​

Async job queue with progress and refund-on-failure · prompt enhancement · credit wallet and per-plan model gating · community publish/share/remix · per-model server-side instructions and token budgets.

2. Roadmap ​

Dated from 2026-08-12. Phases 1–2 are the ones that move the capability numbers, and both are gated on payment details rather than engineering.

Phase 1 — Unlock what is already built (week of Aug 12–19) ​

StepWorkOwnerEffort
1Open self-serve accounts: fal.ai, Runware, ReplicateBusiness1 day, card only
2Add credentials in admin (Providers → Add account)Ops1 hour
3Correct the upstream model ids on every unverified rowEng1 day
4Run the verify sweep, model by model, and read the resultsEng + Ops1 day, ~BDT 500 in real calls
5Retire or fix the 6 failing rowsEng0.5 day

Outcome: 29 → ~60 live models, 20 of them video. No new adapter code.

Step 3 is the real engineering: the unverified rows carry placeholder upstream ids (fal/veo-3.1 resolves to veo-3.1, which is not a fal endpoint path). Each needs the vendor's exact id, which is why verification has to follow.

Phase 2 — Video depth and correct pricing (Aug 19–31) ​

StepWorkEffort
6Price every newly live model from the vendor's real rate card1 day
7Add a tier field to models so Flagship/Premium/Standard is API data, not a document0.5 day
8Map tiers onto plans — Flagship on Pro and above, Standard on free0.5 day
9Storefront: tier badges and per-model sample galleryClient team

Step 6 matters commercially: the current catalogue prices are estimates, and video is where a wrong number is expensive — a flagship clip is roughly 20× a standard image in provider cost.

Phase 3 — Direct vendor accounts, if volume justifies (Sep) ​

Aggregators cost roughly 10–25% more per second than a direct contract and sometimes lag a vendor's newest model by weeks. Once video volume is real, go direct on the two or three models that carry it — Kling (direct also unlocks the lip-sync and extend paths we have already built), Runway, MiniMax/Hailuo. Each is a contract and an invoice relationship: 1–3 weeks of onboarding, no adapter work.

Seedance direct means ByteDance Volcengine ModelArk (CN) or BytePlus (intl), both with business-verification onboarding. Not worth it before volume — take it through fal or Runware first and revisit.

Phase 4 — Auto routing v1 (Sep, 1–2 weeks) ​

See the next section.

Phase 5 — Avatar, music, STT (Sep–Oct) ​

Tavus avatar video and MiniMax music are integrated and unkeyed; ElevenLabs music and Scribe are keyed but failing. Together: ~1 week once Phase 1 is done.

What we are deliberately not promising ​

Timelines above are ours to control. What we cannot promise is a specific third-party model on a specific date — vendors ship, deprecate and re-price on their own schedule, and two of the six verification failures are vendors having changed something under us. The commitment worth making to a partner is the capability: any model reachable over an HTTP API on one of our seven adapter shapes is a catalogue row and a verify run, not a release.

3. Auto routing — can Kikori.ai configure it? ​

Honest answer: routing exists today at the account layer, not the model layer, and everything that exists is ours to configure.

What is live now ​

Per request the gateway picks which vendor account serves a named model, and it is entirely admin-driven:

  • Weights — split traffic across several keys for one provider.
  • Per-account model allowlists — this key serves only these models.
  • Daily and monthly spend caps per key, rolled at the period boundary.
  • Automatic failover — a rate-limited or exhausted key is marked and the next eligible one takes over mid-request.
  • Health-driven exclusion — a provider with no working credential goes INACTIVE and its models stop being offered rather than failing at generate time.

What does not exist yet ​

An Auto mode that chooses the model for the user. Today the client names a model slug. There is no pseudo-model that reads the prompt and picks Flagship vs Standard.

What we would build (Phase 4, 1–2 weeks) ​

An auto slug per modality (auto/text, auto/image, auto/video) resolving through an admin-editable policy:

  1. Eligibility — the caller's plan and the model's live health.
  2. Intent — prompt length, whether an image is attached, requested duration.
  3. Policy — cheapest-that-qualifies, best-quality-within-budget, or fastest, chosen per plan by an admin.
  4. Budget awareness — spend down the wallet on Standard before Flagship.
  5. Transparency — the response names the model actually used, and the decision is logged so a bad route can be explained.

Fully configurable by us, and by design not a black box: same admin surface as the per-model instructions, no client release to change a routing rule.

4. Where the video models come from ​

Ranked by how fast they deliver a live model, with the adapter status we already have.

SourceAdapterGets usSignupCost note
fal.ai✅ doneSeedance, Kling, Veo, Wan, Hailuo and the rest behind one keySelf-serve, card, same dayPer-second, no minimum
Runware✅ doneSeedance, Kling, Wan, Hailuo — usually the cheapest per secondSelf-serve, same dayPrepaid credits
Replicate✅ doneThe long tail, including open-weight videoSelf-serve, same dayPer-second; slower cold starts
Kling direct✅ done (+ lip-sync, extend)Best Kling rate limits and the features we already builtContract, 1–3 weeksCheaper at volume
Runway direct✅ doneGen-4.5, Aleph, Act-TwoContract, 1–2 weeksCredit packs
Luma / MiniMax direct✅ doneRay 3.2, HailuoContract, 1–2 weeks—
ByteDance (Volcengine / BytePlus)❌ not writtenSeedance at sourceBusiness verification, 3–6 weeksCheapest at real volume

Recommendation: open fal.ai and Runware this week. Between them the entire flagship video line — Seedance included — goes live without an engineer writing a new adapter, and having both gives failover when one has a bad day. Keep Replicate as the long-tail escape hatch, and revisit direct contracts once we know which two models the traffic actually lands on.

One caveat to carry into every one of these conversations: model names, ids and prices in this document are what our catalogue asserts, and the video market re-prices monthly. Read the vendor's current model page at signup and let the verify sweep — not the catalogue — decide what we tell a customer works.

kikori.ai — internal documentation