Blog

Guides & tutorials

Practical, no-fluff guides for building with DeepSeek, Qwen, Kimi, GLM and more — through one OpenAI-compatible API.

QwenReleaseAlibaba

Qwen3.8-Max is official: 2.4T parameters, $2/$6 — and the first Max-class open weights

Alibaba shipped Qwen3.8-Max on August 3, 2026: a 2.4-trillion-parameter MoE with 95B active, a 1M context, $2/$6 pricing — and a promise that has never been made before: the Max-tier weights go open next week, alongside a 27B variant. The benchmarks, the caveats, and what two weeks of preview skepticism got right.

Read more
MiniMaxVideo GenerationOpen Source

MiniMax open-sources H3: a 33B video model with native audio — unless you're in the US or EU

MiniMax released the H3 video model's weights on August 3, 2026: 33B dense, 4–15s clips, up to 2K with native stereo audio, running on consumer GPUs via ComfyUI. The catch is in the license — it excludes the US, EU, UK and South Korea. What's actually in the release, what runs locally, and when hosted APIs still win.

Read more
DeepSeekV4 FlashRelease

DeepSeek V4 Flash goes stable: 284B parameters, punching at Opus weight

The stable V4-Flash-0731 landed July 31 — same 284B/13B architecture, retrained, and suddenly scoring 25.2 on Agents' Last Exam against Opus 4.8's 25.7, at $0.14/$0.28 per million tokens. Weights on Hugging Face the same day. The numbers, the post-training story, and what it means in price-war week.

Read more
HappyHorseVideoTutorial

Generate video via API with HappyHorse: the full async workflow

HappyHorse generates 1080p video with synced audio from a prompt, an image, or reference shots — from $0.14 per second. The API is a two-step async task: create, then poll. Complete tutorial with curl and Python, all four models, and the billing rules that matter.

Read more
QwenTutorialAccess

How to use the Qwen API outside China (Qwen3.7-Max and Plus, OpenAI-compatible)

Qwen3.7-Max is Alibaba's shipping flagship tier and Plus its budget sibling — both callable from outside China through an OpenAI-compatible endpoint, international card, no Chinese phone number. Setup, model choice, code, and where Qwen3.8 fits.

Read more
KimiK3Open WeightsSelf-hosting

Kimi K3's weights are out, on deadline. The honest guide to actually using them

Moonshot shipped K3's weights on July 27 — 1.56TB on Hugging Face, exactly as committed. Then come the catches: a new license that gates MaaS resellers, hardware demands that start at 80 GPUs, and a thin serving market full of silent quantization. What is free, what is gated, and what to actually do.

Read more
PolicyKimiQwenTracker

Will the US ban Chinese AI models? What is actually on the table

No ban exists today — what exists is a Treasury review with sanctions explicitly threatened, and the biggest pushback coalition US tech has assembled this year: nearly 200 startups, Nvidia, Microsoft, Meta, Altman and Pichai. The facts, the scenarios, and what to do about your stack, tracked and updated.

Read more
Opus 5KimiGLMPricing

Claude Opus 5 lands at $5/$25 and tops the index. Does the Chinese-model value case survive?

Opus 5 is the new Artificial Analysis leader (61, edging Fable 5's 60) at 26% lower cost per task — the strongest attack yet on the "Chinese models win on value" thesis. We sell that thesis, so here is the honest math: where Opus 5 genuinely closes the gap, and where a 5.7× output-price difference still decides.

Read more
KimiK3Availability

Kimi K3 "sold out": what actually happened, and how to still get access

Kimi K3's subscriptions vanished 48 hours after launch: plans paused, Kimi Code marked sold out, and even paying API users now hit 429 walls. Here is the full picture — why demand broke Moonshot's capacity, what works today, and the practical routes to K3-class capability while it recovers.

Read more
GLMRelease TrackerZ.ai

GLM-5.5: an "epic" tease, a trillion-parameter target, and a gigawatt of Chinese silicon

Zhipu's founder has confirmed GLM-5.5 with an "epic-plus" upgrade promised, an August window reported, and expectations above a trillion parameters — backed by a newly completed 1GW data center running entirely on Chinese chips. What is confirmed, what is expected, and the one question that matters for developers: does GLM stay cheap?

Read more
QwenRelease TrackerAlibaba

Qwen3.8 tracker: the 2.4T preview, the "second only to Fable 5" claim — and the August 3 answer sheet

Alibaba previewed Qwen3.8-Max on July 19, 2026 with a frontier claim and zero published benchmarks. On August 3 the official release landed with a full slate, $2/$6 pricing and a dated open-weights commitment. This page tracked every claim through the preview weeks — here is how each one scored.

Read more
KimiK3BenchmarksPricing

Kimi K3 is out: a 2.8T-parameter flagship, priced like it knows it

Moonshot released Kimi K3 on July 16, 2026: a 2.8-trillion-parameter MoE with a 1M context, the strongest open-weight GPQA score published, and third place on Artificial Analysis behind only Fable 5 and GPT-5.6 Sol. At $3/$15 it is the priciest Chinese-lab model yet. Full breakdown: specs, benchmarks, real cost, and when to pick it.

Read more
KimiK3Release Tracker

Kimi K3 launch tracker: what the leaks got right — and wrong

Kimi K3 launched on July 16, 2026, one day after Moonshot's teaser. We tracked every pre-launch claim — Kivine on Arena, the 2.5T leak, the WAIC timing bets — and here is the scoreboard against the real release, plus where the weights stand. Full specs and pricing in our launch deep dive.

Read more
DeepSeekV4Release Tracker

DeepSeek V4 stable release: the July 24 deadline, peak-pricing reports, and what's rumor

DeepSeek V4 has run as a preview since April 24, 2026. Now reports point to a mid-July stable release with doubled peak-hour API prices — and one hard, official fact: the legacy deepseek-chat and deepseek-reasoner aliases die on July 24. What's confirmed, what's reported, what it means for overseas developers.

Read more
GPT-5.6GLMDeepSeekPricing

GPT-5.6 vs GLM-5.2 and DeepSeek: the price gap that didn't close

GPT-5.6 tier-by-tier against the Chinese lineup, updated for the July 30 price cut: Sol $5/$30, Terra $2/$12, Luna slashed 80% to $0.20/$1.20 — now undercutting DeepSeek V4-Pro. What the cut changes, what it doesn't (flagship coding still costs 4.8× GLM-5.2), and the benchmark caveats.

Read more
GLMOpenRouterUse Cases

What real apps run on GLM-5.2: OpenRouter's top apps by token usage

The apps sending the most tokens to GLM-5.2 on OpenRouter are Hermes Agent, Claude Code, pi, Kilo Code and OpenClaw — mostly open-source coding and agent tools that let users pick any model. Here is what that reveals, and how to run GLM-5.2 yourself.

Read more
ClineTutorialGLMDeepSeek

How to use GLM-5.2 and DeepSeek in Cline (OpenAI-compatible setup)

Cline accepts any OpenAI-compatible endpoint, so you can run GLM-5.2, DeepSeek V4-Pro or V4-Flash inside it through Turiloop — three fields, no code changes, international card, no Chinese phone number.

Read more
Roo CodeTutorialDeepSeekGLM

Connect DeepSeek and GLM-5.2 to Roo Code (OpenAI-compatible)

Roo Code supports any OpenAI-compatible provider, so you can route between GLM-5.2, DeepSeek V4-Pro and V4-Flash from one Turiloop key — cheap by default, escalate when a task is hard.

Read more
ContinueTutorialGLMDeepSeek

Using GLM-5.2 and DeepSeek in Continue.dev (config.yaml guide)

Continue.dev's OpenAI-compatible provider lets you add GLM-5.2 and DeepSeek to VS Code or JetBrains in a few lines of config.yaml — chat, autocomplete and edit, all from one Turiloop key.

Read more
aiderTutorialDeepSeekGLM

aider with DeepSeek and GLM-5.2: the OpenAI-compatible setup

aider connects to any OpenAI-compatible endpoint. Point it at Turiloop with two env vars and the openai/ model prefix to pair-program with GLM-5.2 or DeepSeek from your terminal — international card, no Chinese phone number.

Read more
PricingGLMDeepSeekKimiGuide

Chinese AI model API pricing in 2026: GLM-5.2, DeepSeek, Kimi and MiniMax compared

Official API prices for every major Chinese model, updated July 2026 — Kimi K3 ($3/$15), GLM-5.2 ($1.40/$4.40), DeepSeek V4-Pro ($0.435/$0.87), Kimi K2.6, MiniMax — with the benchmarks that justify them and the closed-model prices they undercut.

Read more
CodingGLMDeepSeekComparison

What is the best LLM for coding in 2026? The honest answer, by budget

Short answer, updated July 2026: Fable 5 still tops SWE-bench Pro, Claude Opus 5 is the new frontier seat at half its price, GLM-5.2 remains the value default, Kimi K3 owns the agentic and front-end ceiling, DeepSeek V4-Pro takes algorithms and V4-Flash the bulk. The benchmarks and prices behind every pick.

Read more
DeepSeekPricingGuide

DeepSeek API pricing in 2026: what it actually costs, and how to pay from abroad

DeepSeek V4-Pro costs $0.435/$0.87 per million tokens, V4-Flash $0.14/$0.28 with cache hits near $0.014 — official June 2026 rates. What drives the real bill, how it compares to GPT-5.5 and Claude, and payment options outside China.

Read more
GLMCodingBenchmarkOpen Source

GLM-5.2 deep dive: the open-weight model that beats GPT-5.5 on coding

GLM-5.2 is the #1 open-weight model and #4 overall, scores 62.1 on SWE-bench Pro, runs a 1M-token context, and undercuts the closed frontier. A technical breakdown of the benchmarks, the MoE architecture, pricing, and how to call it.

Read more
GLMComparisonPricingCoding

GLM-5.2 vs Claude Fable 5: the value champion against the ceiling

Fable 5 is the most capable model money can buy right now. GLM-5.2 is the best open-weight model right now. They cost wildly different amounts. Here's how to decide between them.

Read more
HappyHorseVideoNews

HappyHorse: Alibaba's fast, audio-native video model, live on Turiloop

HappyHorse generates 1080p video with synced audio in seconds, tops the Artificial Analysis video board, and costs a fraction of the field. It is live on Turiloop now — text-to-video, image-to-video, reference-to-video and video editing, from $0.14 per second.

Read more
GuideDeepSeekGLMComparison

Chinese AI models ranked, July 2026: Kimi K3 vs GLM-5.2 vs DeepSeek V4 — what to use for what

Kimi K3 is the new capability king (#3 on Artificial Analysis, behind only Fable 5 and GPT-5.6 Sol). GLM-5.2 is still the value default at $1.40/$4.40. DeepSeek V4 owns algorithms and cheap volume. One page, every pick justified with benchmarks and prices — updated July 25, 2026.

Read more
GuideDeepSeekComparison

How to choose a Chinese LLM API relay (2026 buyer's guide)

GLM-5.2, DeepSeek V4, Kimi K2.6 and MiniMax are excellent and cheap — but reaching them from outside China is the hard part. A practical guide to picking a relay that won't burn you.

Read more
DeepSeekComparisonPricing

DeepSeek V4 vs GPT-5.5 vs Claude Opus 4.8: price and performance (2026)

DeepSeek V4-Pro matches the closed frontier on coding benchmarks at a fraction of the output price. A data-driven 2026 comparison with real benchmarks and API rates.

Read more
GLMKimiComparison

GLM-5.1 vs Kimi K2.6 for coding: which should you use?

Both shipped in the April 2026 Chinese wave and both code well. The real difference is price, context length and agentic strength. Here's how to choose.

Read more
MigrationDeepSeekOpenAI-compatible

Migrating from OpenAI to DeepSeek: a painless, OpenAI-compatible switch

GPT-5.5 output costs $30/M; DeepSeek V4-Pro is a fraction of that and matches it on coding. Here's how to migrate in two lines, no code rewrite.

Read more
DeepSeekReasoningGuide

DeepSeek R1 and its successor: when a reasoning model beats a general one

DeepSeek R1 made chain-of-thought reasoning famous; today that job belongs to DeepSeek V4-Pro's thinking mode, and the deepseek-reasoner alias retires on July 24, 2026. When deliberate reasoning is worth paying for, when it isn't, and how to call it now.

Read more
DeepSeekPricingGuide

DeepSeek V4-Flash: the cheapest capable model in 2026 (and when to use it)

At about $0.14 / $0.28 per million tokens — and $0.014 on cache hits — DeepSeek V4-Flash is roughly 100× cheaper than the closed frontier. Here's what it's good at and when to reach for the Pro tier instead.

Read more
PricingOptimizationDeepSeek

How prompt caching cuts your LLM bill in half (DeepSeek & Kimi cache pricing)

Cache hits cost roughly a tenth of normal input tokens: DeepSeek reads cache at ~$0.07/M, Kimi K2.6 at ~$0.16/M. Structure your prompts right and the savings are automatic.

Read more
TutorialAgentsAPI

Function calling and tool use with DeepSeek and GLM

Tool calling turns a chat model into an agent that can hit your APIs. Here's how to do it with DeepSeek V4 and GLM-5.1 through the standard OpenAI tools interface.

Read more
TutorialRAGDeepSeek

DeepSeek RAG: build a retrieval-augmented pipeline (OpenAI-compatible)

A practical DeepSeek RAG tutorial: retrieve from your own documents and let a cheap, capable model answer grounded in them. A minimal, production-shaped pipeline with the generation code, cost tricks, and an FAQ.

Read more
TutorialImagesAPI

Generating images via API with gpt-image-2

gpt-image-2 produces high-resolution images from a prompt through an OpenAI-compatible endpoint. Here's how to call it — generation and editing — with per-image pricing.

Read more
ProductionReliabilityAPI

Production reliability for LLM APIs: rate limits, retries and failover

A prototype that calls an LLM once is easy. Production traffic needs retries with backoff, timeouts, and a fallback model. Here are the patterns that keep your app up.

Read more
KimiGLMGuide

How to access Kimi K2.6 and GLM-5.1 from outside China

Kimi K2.6 (long-context agents) and GLM-5.1 (coding) are two of the strongest models from the April 2026 Chinese wave. Here's how to call both from abroad, with code.

Read more
TutorialAPIOpenAI-compatible

One key, every Chinese model: an OpenAI-compatible integration walkthrough

A hands-on guide to calling GLM-5.2, DeepSeek V4, Kimi K2.6 and MiniMax through one OpenAI-compatible endpoint — switching models, streaming, and smart routing in real code.

Read more
DeepSeekAPIGuide

How to access the DeepSeek API outside China (no Chinese phone, pay with card)

A practical guide for developers outside China to call DeepSeek's API with an international credit card — no Chinese phone number, no Alipay, OpenAI-compatible in minutes.

Read more