Blog
Practical, no-fluff guides for building with DeepSeek, Qwen, Kimi, GLM and more — through one OpenAI-compatible API.
Alibaba shipped Qwen3.8-Max on August 3, 2026: a 2.4-trillion-parameter MoE with 95B active, a 1M context, $2/$6 pricing — and a promise that has never been made before: the Max-tier weights go open next week, alongside a 27B variant. The benchmarks, the caveats, and what two weeks of preview skepticism got right.
Read more →MiniMax released the H3 video model's weights on August 3, 2026: 33B dense, 4–15s clips, up to 2K with native stereo audio, running on consumer GPUs via ComfyUI. The catch is in the license — it excludes the US, EU, UK and South Korea. What's actually in the release, what runs locally, and when hosted APIs still win.
Read more →The stable V4-Flash-0731 landed July 31 — same 284B/13B architecture, retrained, and suddenly scoring 25.2 on Agents' Last Exam against Opus 4.8's 25.7, at $0.14/$0.28 per million tokens. Weights on Hugging Face the same day. The numbers, the post-training story, and what it means in price-war week.
Read more →HappyHorse generates 1080p video with synced audio from a prompt, an image, or reference shots — from $0.14 per second. The API is a two-step async task: create, then poll. Complete tutorial with curl and Python, all four models, and the billing rules that matter.
Read more →Qwen3.7-Max is Alibaba's shipping flagship tier and Plus its budget sibling — both callable from outside China through an OpenAI-compatible endpoint, international card, no Chinese phone number. Setup, model choice, code, and where Qwen3.8 fits.
Read more →Moonshot shipped K3's weights on July 27 — 1.56TB on Hugging Face, exactly as committed. Then come the catches: a new license that gates MaaS resellers, hardware demands that start at 80 GPUs, and a thin serving market full of silent quantization. What is free, what is gated, and what to actually do.
Read more →No ban exists today — what exists is a Treasury review with sanctions explicitly threatened, and the biggest pushback coalition US tech has assembled this year: nearly 200 startups, Nvidia, Microsoft, Meta, Altman and Pichai. The facts, the scenarios, and what to do about your stack, tracked and updated.
Read more →Opus 5 is the new Artificial Analysis leader (61, edging Fable 5's 60) at 26% lower cost per task — the strongest attack yet on the "Chinese models win on value" thesis. We sell that thesis, so here is the honest math: where Opus 5 genuinely closes the gap, and where a 5.7× output-price difference still decides.
Read more →Kimi K3's subscriptions vanished 48 hours after launch: plans paused, Kimi Code marked sold out, and even paying API users now hit 429 walls. Here is the full picture — why demand broke Moonshot's capacity, what works today, and the practical routes to K3-class capability while it recovers.
Read more →Zhipu's founder has confirmed GLM-5.5 with an "epic-plus" upgrade promised, an August window reported, and expectations above a trillion parameters — backed by a newly completed 1GW data center running entirely on Chinese chips. What is confirmed, what is expected, and the one question that matters for developers: does GLM stay cheap?
Read more →Alibaba previewed Qwen3.8-Max on July 19, 2026 with a frontier claim and zero published benchmarks. On August 3 the official release landed with a full slate, $2/$6 pricing and a dated open-weights commitment. This page tracked every claim through the preview weeks — here is how each one scored.
Read more →Moonshot released Kimi K3 on July 16, 2026: a 2.8-trillion-parameter MoE with a 1M context, the strongest open-weight GPQA score published, and third place on Artificial Analysis behind only Fable 5 and GPT-5.6 Sol. At $3/$15 it is the priciest Chinese-lab model yet. Full breakdown: specs, benchmarks, real cost, and when to pick it.
Read more →Kimi K3 launched on July 16, 2026, one day after Moonshot's teaser. We tracked every pre-launch claim — Kivine on Arena, the 2.5T leak, the WAIC timing bets — and here is the scoreboard against the real release, plus where the weights stand. Full specs and pricing in our launch deep dive.
Read more →DeepSeek V4 has run as a preview since April 24, 2026. Now reports point to a mid-July stable release with doubled peak-hour API prices — and one hard, official fact: the legacy deepseek-chat and deepseek-reasoner aliases die on July 24. What's confirmed, what's reported, what it means for overseas developers.
Read more →GPT-5.6 tier-by-tier against the Chinese lineup, updated for the July 30 price cut: Sol $5/$30, Terra $2/$12, Luna slashed 80% to $0.20/$1.20 — now undercutting DeepSeek V4-Pro. What the cut changes, what it doesn't (flagship coding still costs 4.8× GLM-5.2), and the benchmark caveats.
Read more →The apps sending the most tokens to GLM-5.2 on OpenRouter are Hermes Agent, Claude Code, pi, Kilo Code and OpenClaw — mostly open-source coding and agent tools that let users pick any model. Here is what that reveals, and how to run GLM-5.2 yourself.
Read more →Cline accepts any OpenAI-compatible endpoint, so you can run GLM-5.2, DeepSeek V4-Pro or V4-Flash inside it through Turiloop — three fields, no code changes, international card, no Chinese phone number.
Read more →Roo Code supports any OpenAI-compatible provider, so you can route between GLM-5.2, DeepSeek V4-Pro and V4-Flash from one Turiloop key — cheap by default, escalate when a task is hard.
Read more →Continue.dev's OpenAI-compatible provider lets you add GLM-5.2 and DeepSeek to VS Code or JetBrains in a few lines of config.yaml — chat, autocomplete and edit, all from one Turiloop key.
Read more →aider connects to any OpenAI-compatible endpoint. Point it at Turiloop with two env vars and the openai/ model prefix to pair-program with GLM-5.2 or DeepSeek from your terminal — international card, no Chinese phone number.
Read more →Official API prices for every major Chinese model, updated July 2026 — Kimi K3 ($3/$15), GLM-5.2 ($1.40/$4.40), DeepSeek V4-Pro ($0.435/$0.87), Kimi K2.6, MiniMax — with the benchmarks that justify them and the closed-model prices they undercut.
Read more →Short answer, updated July 2026: Fable 5 still tops SWE-bench Pro, Claude Opus 5 is the new frontier seat at half its price, GLM-5.2 remains the value default, Kimi K3 owns the agentic and front-end ceiling, DeepSeek V4-Pro takes algorithms and V4-Flash the bulk. The benchmarks and prices behind every pick.
Read more →DeepSeek V4-Pro costs $0.435/$0.87 per million tokens, V4-Flash $0.14/$0.28 with cache hits near $0.014 — official June 2026 rates. What drives the real bill, how it compares to GPT-5.5 and Claude, and payment options outside China.
Read more →GLM-5.2 is the #1 open-weight model and #4 overall, scores 62.1 on SWE-bench Pro, runs a 1M-token context, and undercuts the closed frontier. A technical breakdown of the benchmarks, the MoE architecture, pricing, and how to call it.
Read more →Fable 5 is the most capable model money can buy right now. GLM-5.2 is the best open-weight model right now. They cost wildly different amounts. Here's how to decide between them.
Read more →HappyHorse generates 1080p video with synced audio in seconds, tops the Artificial Analysis video board, and costs a fraction of the field. It is live on Turiloop now — text-to-video, image-to-video, reference-to-video and video editing, from $0.14 per second.
Read more →Kimi K3 is the new capability king (#3 on Artificial Analysis, behind only Fable 5 and GPT-5.6 Sol). GLM-5.2 is still the value default at $1.40/$4.40. DeepSeek V4 owns algorithms and cheap volume. One page, every pick justified with benchmarks and prices — updated July 25, 2026.
Read more →GLM-5.2, DeepSeek V4, Kimi K2.6 and MiniMax are excellent and cheap — but reaching them from outside China is the hard part. A practical guide to picking a relay that won't burn you.
Read more →DeepSeek V4-Pro matches the closed frontier on coding benchmarks at a fraction of the output price. A data-driven 2026 comparison with real benchmarks and API rates.
Read more →Both shipped in the April 2026 Chinese wave and both code well. The real difference is price, context length and agentic strength. Here's how to choose.
Read more →GPT-5.5 output costs $30/M; DeepSeek V4-Pro is a fraction of that and matches it on coding. Here's how to migrate in two lines, no code rewrite.
Read more →DeepSeek R1 made chain-of-thought reasoning famous; today that job belongs to DeepSeek V4-Pro's thinking mode, and the deepseek-reasoner alias retires on July 24, 2026. When deliberate reasoning is worth paying for, when it isn't, and how to call it now.
Read more →At about $0.14 / $0.28 per million tokens — and $0.014 on cache hits — DeepSeek V4-Flash is roughly 100× cheaper than the closed frontier. Here's what it's good at and when to reach for the Pro tier instead.
Read more →Cache hits cost roughly a tenth of normal input tokens: DeepSeek reads cache at ~$0.07/M, Kimi K2.6 at ~$0.16/M. Structure your prompts right and the savings are automatic.
Read more →Tool calling turns a chat model into an agent that can hit your APIs. Here's how to do it with DeepSeek V4 and GLM-5.1 through the standard OpenAI tools interface.
Read more →A practical DeepSeek RAG tutorial: retrieve from your own documents and let a cheap, capable model answer grounded in them. A minimal, production-shaped pipeline with the generation code, cost tricks, and an FAQ.
Read more →gpt-image-2 produces high-resolution images from a prompt through an OpenAI-compatible endpoint. Here's how to call it — generation and editing — with per-image pricing.
Read more →A prototype that calls an LLM once is easy. Production traffic needs retries with backoff, timeouts, and a fallback model. Here are the patterns that keep your app up.
Read more →Kimi K2.6 (long-context agents) and GLM-5.1 (coding) are two of the strongest models from the April 2026 Chinese wave. Here's how to call both from abroad, with code.
Read more →A hands-on guide to calling GLM-5.2, DeepSeek V4, Kimi K2.6 and MiniMax through one OpenAI-compatible endpoint — switching models, streaming, and smart routing in real code.
Read more →A practical guide for developers outside China to call DeepSeek's API with an international credit card — no Chinese phone number, no Alipay, OpenAI-compatible in minutes.
Read more →