A reseller will sell you Claude or GPT tokens at 90% off list. [1]The obvious question: how do you undercut Anthropic or OpenAI by an order of magnitude and still make money?
The short answer: someone else pays.
This piece is my own overview of that market. If you've ever wondered where stolen LiteLLM keys end up, how Chinese labs get reasoning traces despite geoblocking, or how a 90%-off token is even possible, this is for you.
Like any market, this one has three parts: supply (where the tokens come from), an intermediary (the gateway), and demand (who buys, and why).

Demand: who's buying, and why
These aren't one kind of customer, but four:
The budget builder. Agents burn tokens faster than ever, and not everyone wants to pay list. A 90% discount is a 90% discount.
The geoblocked user. China and a handful of other countries are cut off by the major labs, partly by political obligation. Providers lean on geoblocking, phone and card verification, and Anthropic has reportedly gone as far as live selfies and government ID for some customers. [2] If you're on the wrong side of that wall and you want frontier models, the secondary market is the door.
The distiller. In February, Anthropic accused Chinese developers of training on its models' outputs, a technique called distillation. [3](A little ironic, given how frontier models got their training data in the first place, but that's a different post.) To collect those reasoning traces at scale you can't just open an enterprise account, you need volume access that doesn't trip the alarms.
The attacker. The newest and most worrying segment. Sysdig documented what it calls a first: a threat actor using an exposed Ollama instance to run an offensive, agentic cyberattack. Stolen AI compute becomes the engine for the next attack. [4]
Supply: where the tokens come from (and why they're so cheap)
Now to the question of the discount raised. If a reseller can charge 10% of list, it's because their own cost is close to zero. There are three main ways to get there:
1. Stolen credentials
Over the last few months there's been a steady drip of AI key leaks. The Cloud Security Alliance's AI safety initiative describes three main vectors: [5]
Hijacking open models. Ollama makes it easy to self-host a model. It also makes it easy to do so badly: Sysdig found roughly 175,000 publicly exposed, unprotected Ollama servers. Free compute sitting on the open internet.
Exploited serving software. A bug nicknamed "bleeding llama" (CVE-2026-7482) [6] let attackers exfiltrate API keys and environment variables from Ollama deployments.
Compromised gateways. LiteLLM, one of the most widely used gateway libraries (around 95 million monthly downloads), was compromised. [7]Gateways are a great target precisely because they hold API keys, and this one reportedly touched more than 2,500 companies.
Outside the CSA's taxonomy, there's a fourth, more improvised route:
Misused customer service chatbots. There were a few instances where public customer service bots could be convinced to generate output that had nothing to do with their core duties. In one case, a GitHub user turned Chipotle's bot into an API. [15] There are also reports of an Amazon bot behaving the same way. [14]Fascinating as this is, I don't think it scales.
The term for this is LLMjacking: using someone else's LLM access without authorization. Stolen keys, hijacked servers, that family of things.
2. Farmed accounts
If you can't steal an existing account, you manufacture new ones. ChinaTalk's reporting [2] describes resellers using Latin American, African, and artificially generated identities to get past sign-up verification. Anthropic's own writing on distillation points the same way: some of the harvested access came from accounts created with fake credentials. Because these accounts run through normal billing, the tokens cost roughly what tokens normally cost, so the discounts here are thinner than at the stolen-credential end. You're paying for laundering, not for theft.
3. Pooled and arbitraged quota
Seat-based plans (Anthropic's $20 / $100 / $200 tiers, say) come with quota many people never fully use. Tools like cli-proxy-api [8]expose the Claude Code CLI as a programmatic API, which means someone can pool unused quota across seats and resell it through a gateway. Add in free start credits and educational accounts. This is the grey end: a ToS violation dressed up as arbitrage.
The size of the discount is basically a function of how the tokens were sourced. Stolen keys mean near-zero cost and the deepest discounts. Farmed accounts mean normal cost plus overhead and modest discounts. Arbitraged quota sits in between. The price tag tells you something about the crime.

The gateway: the intermediary
The gateway (in China often a "transfer station") is the intermediary. It sits between the token source and you, forwarding your requests and passing back the generations. They bundle multiple input keys, route the traffic and enable you to generate new keys. In this way, you won’t even notice if a stolen key gets deactivated, because the gateway shift your workload to the next key automatically.
Infrawatch analyzed around 73,000 gateway servers [9]and found the majority running open-source software, about 40.8% on "new-api" [10]and 10.7% on "cli-proxy-api." Many sit in the US, Singapore, South Korea, Japan, and Hong Kong, often on Chinese cloud providers. Most don't have a slick website; a lot of the trade happens on Telegram and Discord. [2]
But here's the thing I'd most want you to take away. The cheap price isn't the only product. It's the bait. Because every one of your requests passes through the intermediary, the gateway has a second, quieter revenue stream, and it can be worth more than the token margin.
Three things a malicious gateway can do while it forwards your traffic:
Read everything. Your prompts pass through in the clear. Anything in them (a secret, an API key, proprietary code, customer data) can be logged and mined.
Inject into your agent. This is the scary one. Because you're often running an agent in something close to "YOLO mode," auto-accepting its commands, a gateway can insert instructions your system never asked for, say a command to fetch and exfiltrate your credentials. Liu et al. tested this against real routers: across 28 paid and 400 free ones, 1 paid and 8 free routers injected malicious code. Small percentages, but the blast radius on a machine running an autonomous agent is enormous. [11]
Sell your traces. Your interactions are training data. There's no hard proof, but reporting suggests some gateways may resell traces for model training. If you're a company, your proprietary workflows could be quietly financing someone else's model. [2]
And there's a fourth twist that bites certain buyers specifically: some gateways just lie about the model (also called shadow APIs). You pay for the newest Claude, you get cheap open-source tokens dressed up as it. Zhang et al. studied exactly this and suspected the served model wasn't the advertised one in more than 45% of cases. This lands hardest on the distiller from earlier: if you're paying premium prices specifically to harvest frontier-model traces and you're fed open-source output instead, you've paid for nothing. [12]

Operation Bizarre Bazaar, documented by Pillar Security, is one of the bests analysis on real life example. Pillar laid honeypots; attackers scanned them (reconnaissance), validated the access, then exploited and monetized it through a marketplace, silver.inc. Recommended to read. [13]

How big is this, really?
To be honest, I don't have clean numbers, and neither does anyone I've read. But the fragments give a rough sense of scale: ~73,000 gateway servers, ~175,000 exposed Ollama instances, one compromised library with ~95 million monthly downloads reaching 2,500+ companies. I don’t know anything, but this isn't a rounding error or a handful of Telegram scammers. It's an ecosystem, with tooling, marketplaces, and a supply chain.
So what?
A few things I'd take from an overview like this.
First, "the resale token market" isn't one thing. It runs along a gradient from mild (reselling quota you paid for) to criminal (stolen keys plus malware injected into the buyer's own agent), and the location, purpose, and legality vary enormously across it. Lumping it all together as "cheap tokens" misses the point.
Second, and this is the part I'd actually watch, the cheap price is almost never the real transaction. When tokens are 90% off, either someone else is paying the bill or you are the product, through your data, your traces, or your compromised agent. The discount is the hook.
I'd be glad for corrections, especially from anyone closer to the security side than I am.
AI Disclosure: The research, structuring and notetaking was done by myself, Claude wrote a draft that I improved iteratively. No links or sources were added by Claude. Image made by ChatGPT.
