Skip to content

How to Use Claude Code with Kimi K3 (and Switch Models by Typing One Word)

Section titled “How to Use Claude Code with Kimi K3 (and Switch Models by Typing One Word)”

A verified, step-by-step guide to running Claude Code against Moonshot AI’s Kimi K3 using Anthropic-compatible environment variables, plus how to extend the same pattern to GLM and any OpenRouter model

Claude Code doesn’t know it’s talking to Anthropic. It doesn’t check. It reads three environment variables, a URL, a token, a model name, and sends its request wherever those variables point. Change the variables, and the entire tool, the file editing, the terminal commands, the multi-step planning, keeps working identically against a different company’s model entirely.

That’s not a workaround or a hack somebody found. Anthropic built Claude Code this way on purpose, and Moonshot AI (the lab behind Kimi K3) built an endpoint specifically to be a drop-in target for it. The result is that you can set up your terminal so that typing claude launches your normal Anthropic account, and typing kimi launches the exact same tool, same interface, same muscle memory, running Moonshot’s model instead. No proxy server, no router app, no third-party software sitting between you and either provider.

This guide walks through the full setup, verified against Moonshot’s and Anthropic’s current documentation rather than copied from a forum post, plus how to extend the same pattern to GLM and any other model you want on standby, and an honest comparison of when each one is actually worth reaching for.

Moonshot AI shipped Kimi K3 on July 16, 2026: a 2.8-trillion-parameter mixture-of-experts model with a 1-million-token context window. It’s priced at $3 per million input tokens and $15 per million output tokens, roughly a third of what Anthropic charges for Claude Opus specifically (Anthropic’s other tiers, Sonnet and Haiku, are already priced closer to or below Kimi K3). On coding benchmarks it lands solidly in the top tier without topping it: it placed among the top three models across six separate coding benchmarks and led on two of them, while trailing the frontier Claude and GPT tiers on the hardest evals. In blind human-preference testing on frontend coding tasks specifically, it beat every leading US model, developers picking its output over the alternatives without knowing which model produced which.

None of that makes it a universal replacement for Claude. It makes it a genuinely useful second option, one that’s worth being one word away from instead of a full reconfiguration away. Worth flagging directly: K3 launched two days before this was written, so those benchmark numbers come from Moonshot’s own launch reporting and early independent testing, not months of production track record. Re-check against independent evals as they accumulate before making it your default for anything high-stakes.

The 1-million-token context window puts Kimi K3 in the same tier as Anthropic’s largest-context Claude models (an extended, opt-in capability on qualifying Claude models, not their default context size), which matters more than it might seem for coding specifically. A model that can hold an entire mid-size codebase in context at once makes fewer mistakes from forgetting what a function three files away actually does, regardless of how strong its reasoning is on any single isolated problem. Moonshot has also been open about the model’s architecture in a way some competitors aren’t: the underlying Kimi Linear attention mechanism, published as research ahead of the model itself, is designed specifically to make that long a context window computationally practical rather than just technically possible. The full model weights were scheduled for public release shortly after launch, which matters if self-hosting or fine-tuning is ever a consideration down the line, most competing frontier models never publish weights at all.

Before the setup steps, it’s worth understanding the mechanism, because it explains both why this works and why certain things (like the model picker inside Claude Code) behave the way they do once you’ve switched providers.

model picker

Three environment variables do all the work:

  • **ANTHROPIC_BASE_URL**: the endpoint Claude Code sends requests to. Any URL that speaks the same request and response format as Anthropic’s API works here, it doesn’t have to be Anthropic’s own servers.
  • **ANTHROPIC_AUTH_TOKEN**: the credential sent with each request. For a third-party provider, this is that provider’s own API key, not an Anthropic key. Mixing these up is the single most common setup mistake.
  • **ANTHROPIC_MODEL**: the model ID included in each request. This has to match an ID the target endpoint actually recognizes; Kimi K3’s ID on Moonshot’s platform is kimi-k3.

Claude Code reads these from a layered configuration system: shell environment variables, project-level settings, user-level settings, and (for enterprise deployments) admin-managed settings that always take precedence. A CLI flag, --settings <path>, lets you load an entirely separate settings file for a single session without touching your permanent configuration at all. That flag is what makes one-word switching possible: each provider gets its own file, and a small shell function decides which one loads before Claude Code even starts.

Create an account on the Kimi API Platform and generate a key. It’ll be a string starting with sk-. Keep it somewhere you can find it, you’ll paste it into a config file in a moment, not into a chat window or a script that leaves it exposed in your shell history.

Claude Code’s default user-level config lives at ~/.claude/settings.json, but nothing requires you to only use that one file. Create a folder to hold one settings file per provider:

mkdir -p ~/.claude/settings

Then create the Kimi-specific file:

~/.claude/settings/kimi-k3.json
{
"env": {
"ANTHROPIC_BASE_URL": "https://api.moonshot.ai/anthropic",
"ANTHROPIC_AUTH_TOKEN": "sk-your-moonshot-key-here",
"ANTHROPIC_MODEL": "kimi-k3",
"ANTHROPIC_SMALL_FAST_MODEL": "kimi-k3",
"ANTHROPIC_DEFAULT_OPUS_MODEL": "kimi-k3",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "kimi-k3",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "kimi-k3"
}
}

Three details here that aren’t obvious and will cost you a debugging session if you miss them:

The base URL has to end in **/anthropic**, not **/v1**. Moonshot serves two API formats at different paths: one OpenAI-compatible, one Anthropic-compatible. Claude Code only understands the Anthropic-format responses, pointing at the wrong path gets you cryptic parsing errors that look like a broken connection rather than a URL typo.

Set all three **ANTHROPIC_DEFAULT_*_MODEL** variables, not just **ANTHROPIC_MODEL**. Claude Code internally routes some of its own background work, smaller sub-tasks it runs while handling your main request, to whatever it considers the “fast” or “small” tier model. If you only set ANTHROPIC_MODEL, those internal calls can still expect an Anthropic model ID and fail partway through an otherwise-working session. This is the single most common cause of “it mostly works but sometimes breaks” reports.

China-region users need a different domain. If you’re calling Moonshot’s China endpoint instead of the international one, the base URL changes to https://api.moonshot.cn/anthropic.

Open ~/.zshrc or ~/.bashrc and add:

kimi() {
claude --settings ~/.claude/settings/kimi-k3.json "$@"
}

Reload your shell:

source ~/.zshrc

Typing kimi now launches Claude Code with every setting from that file applied for the session. Plain claude, typed with no arguments, still launches with your normal Anthropic account untouched, nothing about your default configuration changed. The "$@" at the end forwards any extra arguments you pass, so kimi -p "fix the failing test" behaves exactly like claude -p "fix the failing test" would, just against a different model.

Step 4: verify the connection actually worked

Section titled “Step 4: verify the connection actually worked”

Don’t take it on faith. Inside a kimi session, run /status, it lists every settings source that loaded and shows the active base URL and model. Then send a trivial test message like hi. A normal reply confirms the full round trip works, not just that the file loaded correctly. (More on why /status specifically, and not the /model menu, in the quirks section below, and what to do if the base URL doesn’t show what you expect is covered in troubleshooting.)

A full walkthrough: one developer’s Tuesday

Section titled “A full walkthrough: one developer’s Tuesday”

Here’s what this actually looks like day to day, not as an abstraction but as a real sequence of terminal commands.

9:15 AM, drafting a quick utility script. A one-off script to parse a CSV export doesn’t need the most capable model available, it needs to be fast and cheap. kimi "write a python script that reads sales.csv and outputs monthly totals as json". Kimi K3 handles it in one pass. Total cost for the exchange: a fraction of a cent.

11:30 AM, a gnarly production bug. A race condition in a payment webhook handler that only reproduces under load. This is exactly the kind of task where paying for the strongest available reasoning is worth it. Plain claude, no flags, back to the default Anthropic account and the top-tier model.

2:00 PM, a large multi-file refactor. Renaming a core abstraction across forty files, mechanical but extensive. Kimi K3’s 1-million-token context window means the whole affected portion of the codebase fits in a single session without hitting context limits, and the task is more about consistent execution across many files than deep novel reasoning. kimi "rename the OrderProcessor class to FulfillmentEngine across the entire src/ directory, updating all imports and references".

4:45 PM, reviewing a teammate’s PR for security issues. Back to claude for the highest-confidence pass on anything security-adjacent, where a subtle miss has real consequences.

None of these required opening a settings menu, editing a config file mid-day, or breaking flow to reconfigure anything. The provider choice became as fast as the task itself.

Two behavioral quirks worth knowing before you hit them

Section titled “Two behavioral quirks worth knowing before you hit them”

Thinking mode isn’t automatic on every Kimi variant. Kimi K3 itself reasons by default and works out of the box. But if you switch to the cheaper kimi-k2.7-code variant instead, that model requires thinking mode to be explicitly turned on (press Tab in Claude Code until you see “Thinking on”), or it rejects requests outright with a 400 error, and features like web search silently stop working until you do.

The **/model** menu inside Claude Code won’t show Kimi models, and that’s expected. Claude Code’s built-in model picker is a fixed list of Anthropic aliases; it has no visibility into what’s actually running behind a custom base URL. Whatever you set in ANTHROPIC_MODEL is what’s serving your requests, regardless of what that menu displays. Trust /status for the ground truth, not /model.

The same three-variable trick works for any provider offering an Anthropic-compatible endpoint. As of mid-2026, Zhipu AI’s Z.ai platform is the other major one, serving GLM models at https://api.z.ai/api/anthropic, and it’s notable as the only other provider besides Anthropic itself with a true Anthropic-compatible route (most others need a translation layer in between).

~/.claude/settings/glm.json
{
"env": {
"ANTHROPIC_AUTH_TOKEN": "your-zai-api-key",
"ANTHROPIC_BASE_URL": "https://api.z.ai/api/anthropic",
"API_TIMEOUT_MS": "3000000"
}
}
glm() {
claude --settings ~/.claude/settings/glm.json "$@"
}

Notice this one has no ANTHROPIC_MODEL line. That’s intentional, not an omission: Z.ai’s Coding Plan auto-selects a GLM model based on your subscription tier rather than requiring you to name one, unlike Moonshot’s setup where the model ID is explicit.

The extended timeout isn’t cosmetic. Cross-provider requests pass through an extra translation step and sometimes run on infrastructure with higher latency from outside the provider’s home region, so a generous API_TIMEOUT_MS prevents Claude Code from giving up on a request that’s slow but still working.

For providers that don’t offer a native Anthropic-compatible endpoint, which is most of them, you can still reach hundreds of models by routing through a gateway like OpenRouter or a self-hosted LiteLLM proxy, which translate Anthropic-format requests into whatever format the underlying provider actually expects:

~/.claude/settings/openrouter.json
{
"env": {
"ANTHROPIC_BASE_URL": "https://router.requesty.ai/anthropic",
"ANTHROPIC_AUTH_TOKEN": "your-gateway-key",
"ANTHROPIC_MODEL": "anthropic/claude-opus-4-8"
}
}

Same shell-function pattern, same one-word launch, any model the gateway supports sitting one alias away.

This is where most guides wave their hands. Here’s the honest breakdown, because “just use whichever’s cheapest” and “just use whichever benchmarks best” are both bad advice on their own.

The pattern underneath the table: pick the model, not the tool, per task, and let the tool be the constant. That’s the entire value proposition of this setup, you’re not choosing a favorite, you’re keeping all the options one word apart and picking fresh every time.

Benchmark tables are one thing, but the number that changes behavior is the one on the invoice. Here’s a worked example, using assumed usage rather than a published benchmark: 40 coding requests in a week, averaging 3,000 input tokens and 1,500 output tokens each (a typical mid-size request with some file context attached). Your own numbers will vary, this is meant to show the shape of the comparison, not a number to budget against directly.

pricing

At this volume, the three options land within a few dollars of each other monthly, which is the honest takeaway: for a single developer doing normal day-to-day work, the cost difference alone rarely justifies the switch. The case for multi-provider setup gets stronger fast at higher volume (an agent running unattended for hours, or a team of ten developers instead of one), where a $6-a-month gap becomes a $300-a-month gap, and at that scale, routing routine work to the cheaper model while reserving the expensive one for what actually needs it is a real budget decision, not a rounding error.

A simple decision rule for which alias to type

Section titled “A simple decision rule for which alias to type”

After using this setup for a while, the choice tends to collapse into a short mental checklist rather than a deliberate decision every time:

  • Is this task something you’d trust a capable junior engineer to do unsupervised? Boilerplate, CSV parsing, straightforward refactors, test scaffolding. Reach for Kimi K3 or GLM.
  • Does getting this wrong cost more than the time saved by using a cheaper model? Security review, payment logic, anything touching auth, data migrations. Reach for Claude direct.
  • Is the task mostly about breadth, not depth? Large mechanical changes across many files, where the 1M-token context matters more than the deepest possible reasoning on any single file. Kimi K3’s context window is a genuine advantage here.
  • Are you not sure yet? Start with the cheaper model. Switching to claude mid-session costs you one retyped command, not a lost afternoon.

Yes. ANTHROPIC_BASE_URL, ANTHROPIC_AUTH_TOKEN, and --settings are documented, supported Claude Code configuration, not an undocumented loophole. Nothing about pointing them at a different provider violates how the tool is meant to be used.

One environment worth flagging separately: if you’re behind a corporate proxy or on a managed device, admin-level settings can override anything in your personal config, including these variables, and some corporate networks block outbound traffic to non-approved domains entirely. If kimi fails to connect on a work machine but works fine on a personal one, check with whoever manages your network policy before assuming your settings file is broken.

Your API keys are sitting in plaintext JSON files. Nothing in ~/.claude/settings/*.json is encrypted. On a shared machine, treat these files the same way you’d treat any other credential file, and double-check your .gitignore if you’re using a project-level .claude/settings.json instead of the user-level one, it’s easy to accidentally commit a working API key into a repository this way.

You’re sending your code to a third party either way. Routing through a different model provider means your prompts, file contents, and codebase context travel to that provider’s servers instead of Anthropic’s. For Moonshot and Zhipu specifically, that means data crossing to infrastructure outside the US by default. That’s worth checking against your company’s data-handling policy before pointing a work codebase at either one, a side project and a client’s production repo are different risk calculations.

**/status** still shows Anthropic’s base URL after launching with your alias. Validate the settings file with any JSON linter, a trailing comma or missing quote is the usual cause, and Claude Code skips a file that fails to parse instead of raising a loud error, so the failure is silent.

Requests work, then fail partway through a longer session. Almost always a missing ANTHROPIC_DEFAULT_*_MODEL variable (see Step 2). Add all three, not just ANTHROPIC_MODEL.

A 400 error mentioning “invalid thinking.” See the behavioral quirks section above, this is the kimi-k2.7-code thinking-mode requirement.

Everything’s slow, not broken, just slow. Increase API_TIMEOUT_MS in the settings file. Cross-border API calls add real latency, and Claude Code’s default timeout assumes Anthropic’s own infrastructure.

The same environment variables that power a shell alias work identically inside a CI pipeline, since GitHub Actions, GitLab CI, and most other runners just set environment variables before invoking a command. If you’re running Claude Code as part of an automated PR review or a scheduled maintenance task, you can point specific jobs at specific models the same way:

.github/workflows/pr-review.yml
env:
ANTHROPIC_BASE_URL: ${{ secrets.KIMI_BASE_URL }}
ANTHROPIC_AUTH_TOKEN: ${{ secrets.KIMI_AUTH_TOKEN }}
ANTHROPIC_MODEL: kimi-k3
steps:
- uses: anthropics/claude-code-action@v1
with:
anthropic_api_key: ${{ secrets.KIMI_AUTH_TOKEN }}
prompt: "Review this PR for code quality and obvious bugs."

This is a genuinely useful pattern for high-frequency, low-stakes automation, a first-pass PR linter that runs on every commit doesn’t need the most expensive model available, and running it on Kimi K3 or GLM instead of Claude can cut the recurring cost of that automation substantially without touching the workflow logic at all. Reserve the Anthropic-backed job for the final review gate before merge, where getting it right matters more than getting it cheap.

Does this count as using Claude Code, or am I just using Claude Code’s interface with someone else’s model? Both, honestly. You’re using Anthropic’s tool, its file-editing logic, its tool-calling framework, its terminal integration, with a different model generating the actual responses. Anthropic built the routing this open on purpose, it’s not a gray-area workaround.

Will this break when Claude Code updates? The three core environment variables (ANTHROPIC_BASE_URL, ANTHROPIC_AUTH_TOKEN, ANTHROPIC_MODEL) have been stable across many releases because they’re part of Claude Code’s documented configuration surface, not an internal implementation detail. The --settings flag is equally official. This is different from relying on an undocumented behavior that could vanish in a patch release.

Can I use this with the Claude Code VS Code extension, not just the terminal? Yes, the same environment variables apply through the extension’s own settings interface, though the exact configuration screen differs from editor to editor. The underlying mechanism, an env block pointing at a different base URL, is identical.

What happens if Moonshot’s endpoint goes down? Your kimi alias fails to connect, same as any API outage would. Nothing about your Anthropic setup is affected, since the two run through entirely separate configuration files. This is actually an argument for having more than one alternate provider configured, if Kimi K3 is unreachable, glm is one word away as a fallback.

Is there a performance cost to routing through a third-party endpoint versus calling Anthropic directly? A small one, typically. Cross-provider requests don’t get Anthropic’s own prompt caching optimizations, and depending on your geographic location relative to the provider’s infrastructure, latency can be somewhat higher. For interactive coding sessions where you’re reading the response anyway, this is rarely noticeable. For high-volume automated workloads, it’s worth benchmarking before committing at scale.

Does switching providers mid-project affect my CLAUDE.md files or Claude Code’s memory of the project? No. Those live in your project directory as plain files, not tied to any provider, so they persist and apply identically whether you launched with claude, kimi, or glm.

Do I need separate accounts and billing relationships with every provider I want to try? Yes, for direct integrations like Kimi K3 and GLM, each is its own account and its own invoice. That’s the trade-off against a gateway like OpenRouter, which consolidates billing across many providers into one account at the cost of a small margin on top of each provider’s raw price and an extra network hop. For two or three providers, separate accounts are simple enough to manage directly. Past that, a gateway starts paying for itself in reduced overhead.

“This is a jailbreak or workaround.” It isn’t. Both sides built this deliberately: Anthropic documents ANTHROPIC_BASE_URL and --settings as supported configuration, and Moonshot publishes an official integration guide specifically for Claude Code. You’re using two products exactly as their makers intended them to be used together.

“Cheaper means worse, so why bother.” Cheaper means differently positioned, not strictly worse. Kimi K3 placed among the top three models on six separate coding benchmarks and beat every leading US model in blind human preference tests on frontend coding specifically. The gap that exists is real but narrower than the price difference suggests, which is exactly why routing by task instead of defaulting to the most expensive option makes sense.

“Once I switch, I lose access to Claude.” No switching happens at all in the sense of losing anything. claude and kimi are two separate commands pointing at two separate configurations. Nothing about your Anthropic account, subscription, or settings changes when you type kimi, and nothing about your Moonshot setup is touched when you type claude right after.

“This only works for Kimi specifically.” The mechanism is generic. Any provider offering an Anthropic-compatible endpoint works identically, and providers without one are reachable through a gateway. The Kimi K3 example in this guide is one instance of a pattern that extends to essentially the entire model market.

“My API key is safe because it’s in a config file, not hardcoded in a script.” A settings file is still a plaintext credential on disk. It’s a reasonable place to keep it for personal use, but it deserves the same handling discipline as any other secret: not committed to version control, backed up securely, and rotated if a machine it lived on is ever compromised.

You’re not picking a winner between Claude, Kimi K3, and GLM. That framing is the wrong question, because the honest answer changes by task, by budget, and by month as each lab ships updates. What’s actually worth having is the ability to stop treating “which model” as a decision that requires reopening a config file, and start treating it as a decision as fast as the task itself. The setup itself takes closer to twenty minutes than ten once you count generating keys and confirming the connection actually works, not ten. After that, every model you’ll want for the next year is one word away.

Further reading: Moonshot’s official Claude Code integration guide, Z.ai’s Claude Code setup docs, and Anthropic’s Claude Code settings reference cover the full configuration surface beyond what’s in this walkthrough.