A New Contender in the Small-Model Arena
ANT Group just dropped Ling 3.0 Tiny onto Vercel's AI Gateway, and it's taking the free slot previously held by Ling 3.0 Flash. If you've been watching the MoE (Mixture-of-Experts) wave, this one deserves a closer look.
The headline numbers:
- 7.9B total parameters, but only ~1.3B active per token
- 256K token context window (yes, you read that right)
- Up to 32K output tokens
- Native function calling and prompt caching built in
- Optimized for responsive agents, instruction following, and multi-turn conversation
That parameter-to-activation ratio is the interesting part. You get the knowledge capacity of a ~8B dense model with the inference cost of something closer to a 1-2B model. For anyone running agents at scale, that math matters a lot.
Why MoE Keeps Winning the Efficiency Race
Traditional dense models activate every parameter for every token. MoE architectures route each token through a subset of experts, so you only pay for what you use. Ling 3.0 Tiny's ~1.3B active parameter count means:
- Lower latency per token compared to a dense 8B model
- Cheaper inference — fewer FLOPs per forward pass
- Better throughput for concurrent agent workloads
The trade-off? MoE models can be trickier to fine-tune and the total VRAM footprint is still tied to the full 7.9B parameters. But if you're doing inference-only workloads, it's a clear win.

Getting Started: Two Lines of Code
The integration with the Vercel AI SDK is genuinely minimal. Here's the entire setup:
import { streamText } from 'ai';
// Use the free tier until Aug 14, then switch to the standard name
const result = streamText({
model: 'inclusionai/ling-3.0-tiny-free',
prompt: 'Summarize this thread and draft a reply.',
});
That's it. The AI SDK handles streaming, retries, and failover through AI Gateway automatically.
Model Name Migration
Pay attention to the naming — this trips people up:
| Period | Model Identifier |
|---|---|
| Until Aug 14, 8:00 AM PT | inclusionai/ling-3.0-tiny-free |
| After Aug 14 | inclusionai/ling-3.0-tiny |
If you hardcode the -free suffix, your app will break on the 14th. Use an environment variable.
Coding Agent Setup
Want to use it inside a coding agent? One command:
vercel ai-gateway coding-agents setup
Then select inclusionai/ling-3.0-tiny-free from the interactive picker. The agent will route all completions through Gateway with your existing auth.

What AI Gateway Actually Gives You
This isn't just a proxy. AI Gateway layers several production-grade features on top of raw model access:
- Unified API — one interface for every provider
- Usage and cost tracking — per-key, per-model breakdowns
- Retries and failover — automatic fallback if a provider goes down
- Zero Data Retention support for compliance-sensitive workloads
- Budgets per API key — hard caps so an agent loop can't drain your card
- Routing rules — send cheap requests to small models, complex ones to frontier models
- BYOK (Bring Your Own Key) — use your own provider keys, still no platform fee
Pricing is pass-through. No markup on inference, no platform fee, even on BYOK requests. That's a meaningful departure from most gateway products that skim 10-20% off the top.
Limitations and Caveats
Before you rip out your existing setup, keep these in mind:
- Free tier expires August 14, 8:00 AM PT. After that, it's pay-per-use at provider rates. Budget accordingly.
- MoE ≠ dense quality. A 1.3B active model won't match a 70B dense model on hard reasoning tasks. Great for agents, summarization, and multi-turn chat — not for solving novel math proofs.
- 256K context is impressive, but attention cost scales quadratically. Long-context performance can degrade in ways that benchmarks don't always capture.
- Vendor lock-in via Gateway. Yes, it's a unified API, but migrating away still means rewriting your routing logic.
- Function calling is 'native' — verify it. Always test tool-calling reliability against your specific schema before shipping to production.

Should You Try It?
If you're building agents, prototyping multi-turn chat, or just want a cheap model for summarization pipelines, yes. The free window until August 14 is a low-risk way to benchmark Ling 3.0 Tiny against whatever you're running now.
The bigger story here is the MoE trend. Small-active-parameter models with large total capacity are becoming the default choice for agentic workloads, and Ling 3.0 Tiny is a solid entry in that category.
Next Steps
- Swap the model string in an existing AI SDK project and compare latency
- Test function calling against a real tool schema (not just the playground)
- Set a budget cap on your API key before you deploy — trust me on this one
- If you're exploring on-device inference, check out how Google's AI Edge Gallery is pushing cross-platform function calling — the two trends are converging fast
Related Reading
- How Messenger's Advanced Browsing Protection Works Without Compromising Privacy — a good reminder that privacy and capability aren't always zero-sum.