Aldous Huxley’s book Brave New World is one of those great works of literature that we have all heard about and almost no one has actually read. It is a dystopian novel about how technological progress leads to a different world. With the release of the ChatGPT genie, the world entered an exponential and probabilistic era.
The growth in AI (AI adoption, the usage of ‘tokens’, size of AI models etc) has been so tremendous that they have to be depicted on logarithmic scale in charts. To human minds accustomed to linear processes, this exponential increase is hard to fathom.
Leading up to this ExPro era (exponential and probabilistic), we will encounter many paradoxes. Multiple contradictory things will hold true simultaneously. An unprecedented pace of progress in one domain (Artificial Intelligence) will be constrained by the physical realities of the real world (electricity, supply chains etc.). Market predictions and prophesies may come true but we won’t know which ones and when.
We have already seen a lot of interesting developments that have defied conventional wisdom.
- The presence of Open Source/Open Weight models from China that are practically free were supposed to drive the expensive Frontier Models from Anthropic and Open to extinction. Both coexist.
- Frontier Models are not only thriving but have become pricier. This pricing premium stays even though the Open Source/Weight are almost as good as the Frontier models.
- Nvidia, the arms supplier to everyone in the current AI wars, is also the most vocal supporter in the US of Open Source models. These Open Source models require a fraction of the hardware that latest Frontier models use. Why does Nvidia promote AI models that dampen the demand for its own products?
- China’s Open Source/Weight AI models are critical in democratizing AI.
- Why the strongest player in AI actually looked like the first one to bite the dust.
At some point, we will hit a ceiling because…Scaling Laws. And the nature probabilistic nature of LLMs. At that point, the world would have fundamentally changed. AI would have proliferated. Your spectacles could ‘talk’ to your car with the roads eavesdropping.
Who knows how complex systems with emergent properties react to exponential yet probabilistic change. What we do know is that the world would have fundamentally changed.
Table of Contents
- Introduction: Living in an ExPro World (Exponential and Probable)
- Exponential and Probable dynamics
- Futility of prediction
- AI Bubble fears vs coming AI advances
- The 3 Scaling Laws & Jevons Paradox
- The 4 Waves of AI
- The Lobster Rescues Markets from the Whale
- DeepSeek’s disruption
- OpenClaw and the rise of Agentic AI
- The Third Wave of Agentic AI
- But where is the Money?
- Jevons Paradox and token usage
- Anthropic’s spectacular rise
- Tokenmaxxing and ROI fears
- The Rarest Commodity in the Universe
- Waterloo and Rothschild’s information advantage
- Information vs. Intelligence premiums
- Black–Scholes, Jim Simons, and actionable intelligence
- Mythos: The Agentic Employee Arrives
- Release of Mythos
- Test‑time scaling and inference scaling laws
- Emergence of Agentic Employees
- The Shot Fired at the End of the Race
- Anthropic vs. OpenAI
- Compute shortages and token limits
- Nvidia’s moat and CUDA independence
- The AI Agnostics and Skeptics
- Gary Marcus and fallibility of LLMs
- Agentic AI as “useful AI”
- Scaling laws still intact
- Hype, pessimism, and doomerism
- Open Source Models & China
- Rise of Kimi, Qwen, DeepSeek
- Algorithmic efficiency vs. brute force
- Distillation campaigns and geopolitical angles
- Look East > Frontier Models in China
- Look East > Frontier Hardware in China
- Conclusion
LLMs are amazing but also sometimes suck
AI responses are astonishing, especially when we first encountered them. Yet they don’t have the absolute certainty we have come to expect from computers. The comparison obviously isn’t fair- an Excel formula works on clearly defined input and rules and the output is deterministic. Large Language Models (LLMs), on the other hand, work with hazy inputs (‘What’s the weather like today’) and the output is usually the answer with the most ‘likelihood’.
Just as words are the most instinctual way humans communicate, LLMs’think’ in numbers. If you look inside the brain of an LLM, you can see words of the query being converted to numbers (’embeddings’/’vectors’). These numbers are multiplied with another set of numbers called ‘weights’ that were calculated when the LLM was trained (or ‘taught’) how to think. Probabilities are assigned to responses and the responses are translated back into words – and that’s your answer.
In eschewing the certainty of traditional software that relied on defined inputs and outputs, the LLMs were able to outgrow the shackles of traditional computing systems. They became so bold as to attempt things that previously only human minds could do – understand a problem, apply ‘intelligence’ and generate solutions to open-ended queries.
The probabilistic nature of AI allows it to do wonderful things. The lack of predictability of LLMs that we have come to expect from traditional computing systems however, will periodically lead us back to the same question – is AI overhyped and is the bubble over?
Is the AI bubble over?
In early 2025, a general sentiment existed that AI capabilities had peaked. A better appreciation by the wider public of the probabilistic nature of LLMs led to the magical becoming mundane.
What the public couldn’t see was that the models had become exponentially ‘better’ or more useful. The latest version of ChatGPT is at least 100x better than the original ChatGPT in most parameters – even if it doesn’t seem like that.
Can AI could do anything useful was the next question to prop up. Or was it simply an advanced version of Google search?
After all, ‘Generative AI’ could only come up with words and images – that too, never 100% accurate. However fancy these words and images were, they couldn’t actually do anything. Talks of an ‘AI Bubble’ gained ground. It was a strange bubble where the reigning Champion of AI, Nvidia, was at decadal low valuations but unheralded companies that added ‘AI’ to their name saw their stock price go vertical. Eventually even shoe companies figured out this trick.
To get a sense of where AI was headed, one must heed the words of the Prophet of AI, Jensen Huang. The ‘Profit of AI’ would be a more appropriate epithet given that all the profit of the Generative AI gold rush went almost exclusively to Nvidia in the initial years. This profit concentration strengthened the belief that everyone would lose money buying expensive GPUs from Nvidia but nobody would make money selling AI services due to fierce competition. After all, AI was available for free except the top tier models.
Jensen Huang predictably maintained that we are in the early days of AI. He laid down the 3 Scaling Laws and the 5 Waves of AI to support his claim.
The 3 Scaling Laws
The AI Scaling Law postulates that AI (LLMs) will get better if you increase the Parameters, Training dataset size and the Training budget. So if you exponentially increase the GPUs and model size, the AI will become exponentially better than the previous version.

Jensen Huang categorized these into 3 Laws:
- Pretraining
- Post‑training
- Test‑time scaling (Inference)
The 3rd Scaling Law came in around the time that murmurs were growing about AI being a party trick. Earlier LLMs gave a one-shot response to any question. Test Time Scaling allowed the LLMs to ‘think’ and ‘validate’ their responses.
OpenAI o1 model was released in September 2024 with Chain-of-Thought (CoT) and Self-Correction features that enabled ‘deep thinking’. The magic was back! Just like humans, if you give machines time to think, they can come up with much better responses.
This exponential improvement in AI changed the AI market dynamics. Earlier, most of the AI hardware was geared towards Training. This market was dominated by Nvidia. The new ‘thinking’ and ‘reflecting’ AI models led to a massive demand in Inference (where models already trained provide responses to users’ queries). This new Inference market soon dwarfed Training and a lot of new players began cementing their niches.
The Hyperscalers such as Google, Amazon and Microsoft came up with home-baked chips to reduce dependence on Nvidia for Inference. Smaller nimbler players turned science projects into creative ways to address Inference demand. This included Groq, Cerebras, Samba Nova etc. Don’t fret for Nvidia though. Losing small amounts of market share in an exponentially increasing Inference market still meant ungodly amounts of profits.
The markets recovered but they were soon in for a shock.
The DeepSeek head fake and Jevons Paradox
DeepSeek’s R1 model was released around early 2025. This model was Open Source and developed at a fraction of the cost incurred to Train the latest OpenAI o1 model. In terms of capability, however, it was almost as good as OpenAI’s latest model . Over one eventful weekend, articles on Twitter and Substack anxiously began to confront two conclusions:
- Are Scaling Laws irrelevant? If China can develop AI models without Nvidia’s latest hardware, are the massive investments in data center hardware a waste of money?
- Is AI now commoditised? If China’s model is as good as American models and free, who pays for AI?
In the broader public’s mind, DeepSeek unfairly gets credit for Test‑time scaling (Inference). DeepSeek does get credited, probably fairly, for ‘democratizing’ extended Inference. Their models were all Open Source (more accurately ‘Open Weight’) and could be used freely unlike the western/US companies that charged for frontier AI models.
The legendary investor Peter Lynch once said:
“There is always something to worry about. Avoid weekend thinking and ignoring the latest dire predictions of the newscasters.”
The market opened the next day and promptly lost $1 trillion, with Nvidia itself losing $596.7 billion in market cap. The prognosticators and economists incorrectly surmised that the sky is falling.
In all this doomerism, Groq CEO Jonathan Ross, raked up a 19th century ‘law’ to explain why the market got it all wrong. In fact, the opposite was more likely – the AI economy would grow further. As ‘Intelligence’ became cheaper and better, its usage would grow exponentially.
This was first observed by English economist William Stanley Jevons in 1865. As steam engines became more efficient and used less coal, the overall usage of coal grew instead of declining. More steam engines were built as the costs to run them declined. At some point, it just made sense to run steam engines that used coal vs horses that ‘ran’ on hay.
Back in the 21st century, the markets started recovering on hopes of Jevons Paradox. Satya Nadella and others also joined the chorus and explained why cheaper tokens are good for everyone.
The 4 Waves of AI
Abundant tokens led to increased ChatGPT sessions. AI usage grew exponentially, as laid down by Jevons Paradox. In mid-2025 the markets oscillated to AI-forever optimism from the doomerism triggered by DeepSeek.
As the markets reached new highs, familiar doubts began creeping up again. The ‘bubble’ world started popping up everywhere. The two big fears resurfaced:
- Can AI ‘do’ anything?
- Will anyone make money out of AI except Nvidia and its supply chain (semiconductor industry)?
It is Jensen Huang’s lot in life to explain with patient exasperation a) that AI has just begun and b) the best is yet to come. He had explained about the 3 Scaling Laws and everyone nodded along. He tried another approach.
There are Four Waves of AI and we have barely reached the second wave:
- Wave 1: Perceptive: AI learning to perceive and understand its environment
- Wave 2: Generative: Generating entirely new content (text, images, audio and video)
- Wave 3: Agentic: AI agents execute tasks independently over extended periods of time
- Wave 4: Physical: AI is embodied in the physical world (robots, cars, factories)
The Lobster rescues Markets from the Whale
DeepSeek, symbolised by a blue whale, had shaken the market’s confidence in AI. Another sea-faring creature, the lobster, emerged to send the market back to manic optimism.
In Nov 2025, Austrian developer Peter Steinberger released OpenClaw. It was originally called Clawdbot as a pun/homage to Claude (the Anthropic AI agent) and featured the now famous red lobster.
This moment rivalled the release of ChatGPT three years ago in terms of sheer impact.
For the first time, AI could actually do something. After all, Humanity didn’t invent AI only to be ordered around by it. At first it was all fun and games as hackers and intrepid souls experimented with this Open Source tool. Emboldened by how easily Agents did routine tasks, some trusted it with credit cards. Others gave Agents greater permissions to their computers or cloud. Were some databases deleted? Yes. Did some lucky guy get 200 McNuggets? Yes.
Mishaps aside, developers and companies began to realise the enormity of Agentic AI. A clear trend was visible – an exponential explosion in the use of tokens. Not only were developers using AI, the AI Agents created by them were also using AI.
The Third Wave of Agentic AI had arrived.
…but where is the money?
AI Agents working all night is well and good but is anyone making money using AI? While OpenClaw had addressed the first major ‘bubble’ concern (Can AI do anything useful?), it only deepened the second fear about Return on Investment (ROI). With Agents guzzling up tokens, AI costs would skyrocket. With no corresponding revenue to offset the increased costs, would the AI bubble burst soon?
To recap the story thus far, hardware (Nvidia GPUs, Networking, TSMC process nodes) and software innovations (CUDA, distillation, quantization, MOE & improvements in Nvidia CUDA), had led to a 10x decrease in AI costs, as measured in tokens.

This 10x fall in the cost of tokens has led to a predictable rise in the use of tokens (AI usage). The top 2 users of tokens were Google and OpenAI in early 2026. Both have seen an approximately 10x rise in tokens consumed between Jan 2024 to Mar 2026.

Jevons Paradox did work! The approximate 10x decrease in token costs actually led to a 10x or greater increase in AI usage.
But did anyone make money, except Nvidia?
Well, Anthropic did. In the most spectacular fashion.
The obvious use case of AI was generating code – perfect for LLMs that can effortlessly synthesize text in meaningful patterns required for coding. The rise of Agents meant that code could be written and executed without human bottlenecks. Aided by ‘Tokenmaxxing’, the revenue started rising steeply to the point of going vertical.
Anthropic making tons of money using AI
The question now is whether these revenues are sustainable or a fad enabled by ‘Tokenmaxxing’ trends. A lot of companies just gave targets to their employees to use ‘tokens’ or lose their jobs. Predictably, employees began burning up tokens doing all kinds of activities, some of them actually productive!
Why were the smartest companies in the world giving arbitrary token usage targets to their employees? Didn’t they know that most employees would find ingenious ways of meeting the target without doing anything useful? One of the hopes was that growing adoption of Intelligence could lead to some rare use cases that could have disproportionate results on profits. But the main contributor was probably the fear that competition might find better uses of AI. Using up tokens was as much about opening up economic opportunities as it was about denying the competition any advantages.
We are now seeing increasing evidence of companies throwing up their hands and saying ‘No mas’. CEOs of Snowflake and Shopify have publicly spoken out against Tokenmaxxing. Uber CEO Dara Khosrowshahi had to do a course correction after the company used up all the AI budget for a full year in the space of a few months.
With Tokenmaxxing giving way to common sense, will Anthropic usage continue to go up or is it downhill from here?
The Open Source models from China are doggedly trailing frontier AI models by a few months. These Open Source models are practically free, so does it make sense to pay top dollar for Intelligence that can be achieved for free in a few months?
If enough companies switch to Open Source models, will the AI buildout collapse?
The rarest commodity in the Universe
In June 1815, Napoleon’s Grande Armée was defeated at Waterloo by the combined forces of the British Duke of Wellington and Prussian Field Marshal Blücher.
Back in London, the news of the famous victory hadn’t reached and the mood was sombre. Napoleon had already executed his old playbook – defeat the Prussians in Ligny, engage the British at Quatre Bras and prevent the enemies from combining forces.
The defeated Prussian forces were expected to retreat along their supply lines back to Germany. Napoleon would finally fight and defeat the now isolated Iron Duke Wellington himself.
In this gloomy environment, Baron Nathan Mayer Rothschild began selling ‘Consols’ – bonds issued by the UK government. This confirmed the worst suspicions – “Rothschild knows…” people whispered. “Waterloo is lost.” The value of the Consols plummeted and at the final hours of trading, the Baron scooped up the cheap Consols. News of Napoleon’s unexpected defeat at Waterloo finally reached London. The value of the Consols began rallying sharply. And thus began the Rothschild fortune.
Or so the story goes. While this popular legend may not be entirely accurate, it does highlight the tremendous opportunities for making money if you are privy to information that is not known to the rest of the financial world.
Baron Rothschild had gained an advantage of a few days with the knowledge of the victory in Waterloo. Allegedly, but let’s not let facts get in the way of a good story.
In today’s markets, an information advantage is measured in milliseconds and nanoseconds. Flash Boys by Michael Lewis describes the lengths to which companies can go to get information nanoseconds before their competitors. One such example described in the book is a $300 million project in 2009 to construct an 827-mile (1,331 km) fiber-optic cable ‘that cuts straight through mountains and rivers from Chicago to New Jersey—with the sole objective of reducing the transmission time for data from 17 to 13 milliseconds between the Chicago Mercantile Exchange and Nasdaq.’
AI doesn’t give you an information advantage yet. It offers, possibly, something more valuable – actionable intelligence.
A classic example of actionable intelligence is the Black–Scholes model, developed in 1968. It is used to price options and moves trillions of dollars in the stock market. The Nobel Prize for Economics was awarded to Fischer Black and Myron Scholes in 1997 for formulating the model. Edward Thorp, however, had developed the Black–Scholes formula in 1967 but decided to milk the Options market. He sacrificed academic glory and a Nobel Prize to make serious money for years until others caught on to the Black-Scholes model.
A more eye popping case is Jim Simons, founder of the Medallion Fund. He ‘achieved 66% annual returns for over three decades — outperforming even Warren Buffett for that duration. When he died in 2024, his net worth was estimated to be $31.4 billion. Not bad for a Professor who started a new career as a Hedge Fund manager when he was 40 years old.
Early access to information has always been valuable. To convert this information to opportunity, you need Intelligence. Nowhere is this more evident than financial markets.
As Jeremy Iron’s character explains in the movie Margin Call “Be first, be smarter or cheat”.
If a better AI model allows you to be first or ‘smarter’, it will command a pricing so premium that it seems unreasonable bordering on insane. These ‘Frontier tokens’ will remain in demand, even if there are free AI models that give the same results one second later.
One such model is Anthropic’s Mythos.
Mythos
The release of Mythos was another roller coaster ‘here-we-go-again’ ride. Usually, there is a lull between AI breakthroughs as AI companies wait for the latest hardware from Nvidia. The new GPUs from Nvidia ship every year and are usually 3-4x more capacity than the previous version for training new models. This gives the world an year’s time to become accustomed to the current AI models.
This time however, a staggering development happened immediately after the release of OpenClaw – the release of Mythos model by Anthropic. For the first time, the world began to see a future where AI could realistically replace human employees.
The Mythos model was able to ‘work out problems’ or ‘apply intelligence’ for extended periods of time (Test Time Scaling) and run autonomously (Agentic AI) until they achieve the result. This is what employees in a firm do – work on objectives decided by top management over a defined period to deliver tangible outcomes.
Thus far, AI could only provide information. Now AI could persistently work on problems and accomplish actions for periods up to weeks. Interestingly enough, the Inference Scaling Laws held up. AI models could ‘think’ longer and it improved the results. Just as Jensen Huang predicted! If a model could be assigned tasks independently over periods of one month plus, we could see the emergence of ‘Agentic Employees’.
The Shot fired at the End of the Race
In the AI race, the gun is placed at the finish line and the shot is fired after the race. By the winner. At the runner-up.
No moat built on software is safe now. Google found that out the hard way as it narrowly avoided being disrupted by OpenAI’s ChatGPT. Soon, OpenAI itself was going to be disrupted by Anthropic’s Mythos.
Given the aura of invincibility that Mythos had, enterprises were happy to pay through the nose for access. The momentum had shifted decisively in Anthropic’s favour. OpenAI, Microsoft, Nvidia and Elon Musk watched nervously.
OpenAI’s models competed with Anthropic and it was clear that the latter had better products. The Chinese Open Source/Weight models would clean up the rest of the AI market leaving no place for OpenAI. The chief patron of OpenAI was Microsoft and it faced two challenges – Anthropic was not available on its cloud and its OpenAI investments would be decimated. Nvidia couldn’t justify charging 75% gross margins for its products if Anthropic could simply build its latest models on Amazon/Google’s hardware. Elon’s SpaceX had the opposite problem of Anthropic’s – massive clusters of Nvidia hardware but little demand for its AI model Grok.
To the relief of all four, Mythos was so successful that Anthropic ran out of compute! Dario Amodei had paced himself for the long run as opposed to Sam Altman who spent money ‘like a drunken sailor’ to get access to ‘compute’ (chips, data centers etc). Dario hadn’t anticipated the wild success of Mythos and he just didn’t have the capacity to meet this tremendous demand. Arbitrary limits were imposed on Mythos tokens, causing an uproar. A generation now hooked to the good stuff just couldn’t live without 24×7 access to premium tokens.
With Mythos, Dario was in a position to deliver the coup de grâce to OpenAI. But Dario’s prudence let Sam Altman back into the race. While Anthropic was busy arranging compute, OpenAI released GPT-5.6 within a month. This new model came close to the legendary Mythos but more importantly, was available for use right now . OpenAI stopped the hemorrhaging of its customers and was back in the race.
Anthropic had proven that you didn’t need Nvidia’s hardware to develop the best AI models. Nvidia’s moat was widely attributed to CUDA – the ‘software’ that configured GPUs to do the math necessary for AI. Anthropic had spent considerable engineering resources to ensure it wasn’t dependent on CUDA. But if it had used the brainpower at its disposal on developing the best models quickly, could it have gained an unassailable lead over OpenAI?
Not one to cry over spilt milk, Amodei got into action. Nvidia had all the chips in the world. Elon Musk had a data center full of Nvidia chips but no demand for its Grok AI model (mostly capable with the odd rant here and there). Demand met Supply. Dario finally got access to compute to meet the burgeoning demand for its AI.
The AI race continued as before with OpenAI pitched against Anthropic, the former now having conceded first place. Nvidia showed its chips were essential, even if you worked around their moat. Microsoft got into a new agreement with OpenAI and also allowed access to Anthropic on its cloud. Elon could ‘IPO’ SpaceX at an Elon premium, having demonstrated that no one could put together infrastructure in the ‘real world’ of steel, turbines and batteries like SpaceX.
The Mythos scare reinforced a message – no amount of money spent on AI was excessive. The real message was that no one really knows what’s going to happen. And in this new environment, we need to listen to the AI Sceptics.
The AI Agnostics and Sceptics
A lot of the faith in spending large amounts of capex to design LLMs was based on the hope that eventually we will create AGI – Artificial General Intelligence. If AI can achieve any cognitive task that humans can do, then any amount of current investment is justified.
Gary Marcus, chief AI sceptic of our times, kept proving that LLMs are fallible. They may achieve fantastic feats like conjuring new protein molecules but trip up occasionally when counting the number of ‘i’s in ‘institute’. The LLM approach will clearly not lead to AGI or any AI Deity. The hopes and buzz of AGI disappeared but the capex boom continued in the hope that ‘some day’…
Agentic AI and its Mythos manifestation showed that you don’t need an Einstein to do most activities. We are now on the threshold of Agentic Employees. Jensen Huang, possibly the last human CEO of Nvidia, mentioned in Computex 2026 that there are 30 million developers worldwide who generate $3 trillion in economic output.
“Useful AI has arrived” he declared. The AI Total Addressable Market (TAM) now had a minimum tangible but viable number – $3 trillion.
It is now clear that coding is the most natural use case of LLM-based AI models. Andrej Karpathy, undisputed world champion of coding according to tech and podcast bros, is on record stating that 80% of his code is AI generated. In one month, AI went from generating 20% to 80% of the code he ‘wrote’.
And all of this happened without even the latest hardware from Nvidia! In fact, Anthropic achieved this feat using Amazon’s Trainium chips and some Google TPUs. Jensen Huang, long may he live, said Claude Mythos was trained on a “fairly mundane capacity, and a fairly mundane amount of it”.
And mind you, Training Scaling Laws are still in effect – the new AI models will get better as they get trained on more GPUs. The Inference Scaling Laws are also holding up – an existing AI model can generate better results if given more time to ‘think’. As the models keep getting better exponentially, what further capacities can they unlock?
The world lurches from dismissive pessimism (‘it’s all hype’) to irrational optimism (‘end-of-jobs’) to doomerism (‘we are done for’). With Mythos, all of this happened in a period of a few months.
The model was so powerful that it could not be released publicly. Or so said Dario Amodei. Incidents of cybersecurity vulnerabilities uncovered after decades came out. Mythos was powerful. But, as Gary Marcus pointed out, it was also a bit…what shall we say…hyped.
“Scare, hype, release. And repeat.” was the familiar pattern that Gary observed when companies release new models.
The US Government gladly bought the hype and banned the model to foreign users or entities. Speculation abounded on the real motives behind the ban; the US Administration had an axe to grind against Dario, Anthropic cried wolf/painted a devil on the wall and got its comeuppance or maybe the model was as lethal as advertised.
There remained one more possibility – to prevent Chinese Open Source companies from ‘distilling’ Mythos and closing the Frontier vs Open Source gap. As all the top Open Source AI models are Chinese, the geopolitical angle cannot be ruled out.
‘Open Source’/’Open Weight’ Models & China
All AI benchmarks (MMLU, GPQA, HumanEval, SWE-bench, AIME) are unanimous that the top Open Source models are from China – Kimi (Moonshot AI), Qwen (Alibaba Group) and DeepSeek.
Jensen Huang mentioned that China has around half the AI researchers in the world. So it is entirely natural that five kids with a few GPUs in Hangzhou could come up with the next top Open Source AI model. DeepSeek, the company that rocked the Western financial markets, is primarily a Hedge Fund who happened to develop the world’s best Open Source AI model. DeepSeek’s success not only caught the global financial system off guard, it also surprised Chinese Government who were expecting the national Champions (Alibaba, Baidu) to deliver first.
Non-Chinese models that occasionally make it to the top are LLaMA (Meta), Mistral and Nemotron (Nvidia). They are soon crowded out by newer AI models from China. Companies like Anthropic and OpenAI also offer Open Source flavours of their top models but their heart isn’t in it. Why cannibalise revenues from the top models by offering free versions of the same.
We know that US has banned sale of the most advanced GPUs (from Nvidia, AMD etc.) to China. The US Government had, on paper, allowed the sale of less powerful GPUs to China. However, regulatory flip-flops and bureaucratic opaqueness had practically shut down the sale of even these chips. By creating uncertainty on policy, the US possibly hoped to keep China in limbo, forever waiting for the latest chips. The Chinese government also banned the same watered-down chips, as it looked to create a domestic chip ecosystem independent of US. Or maybe China understood that US had no intention of ever selling even the low-end GPUs to China.
Whatever be the effectiveness of the US ban, two things are clear:
- None of the top AI models are from China but…
- The best Open Source models are all from China and they are only a few months behind the top models.
If Dario Amodei was a Bond supervillain, he could have blackmailed the world with Mythos – it was supposedly that advanced. Well guess what? In a few months, the Chinese Open Source models have caught up.
So how does Zhipu, the latest Open Source kid on the block, catch up with the aptly named Mythos? To simplify, the US/Western models rely on brute force. With Scaling Laws still intact, they keep adding ‘moar’ GPUs, ‘moar’ CPUs, networking etc and the models keep getting better. The Chinese models have adapted to the scarcity in AI hardware by being creative on how they use the limited hardware. They have relied on Algorithmic efficiency and energy abundance. Examples of Algorithmic efficiency are Mixture-of-Experts (MoE) models, activating specific parameters instead of the entire models (for ex, using 20 billion parameters instead of the entire 200 billion) etc.
But is there any other reason why the Open Source models catch up quickly with the top AI models despite the US ban on advanced hardware/GPUs?
The White House and Anthropic don’t usually see eye-to-eye but they agree on one thing – ‘Chinese entities were running “industrial-scale” distillation campaigns against American frontier AI models, leveraging “tens of thousands of proxy accounts” to evade detection.’ (Source)
Whatever be the reasons, we can assume that in a few months, Open Source AI models can match the capabilities of Frontier Models for a fraction of the cost. Even the mighty Mythos had to share the podium in just a couple of months with the latest upstart model from China.
And so the questions came back to haunt us. Can Anthropic/OpenAI continue to charge exorbitantly for tokens when the Chinese models are just as effective?
And another question also comes up – Do we need Nvidia at all? Mythos was ‘developed/’Trained’ largely on Amazon’s custom chips (Trainiums) and not Nvidia. And Open Source models can run on your local computer stacks of Apple Mac mini / Mac Studio – again no Nvidia needed.
We will find the answer in the most unexpected place – China, where Nvidia chips are banned. But before that, let’s understand the Open Source landscape.
Open Source Models & America
Microsoft has recently thrown its weight behind Open Source.
Satya Nadella, CEO of Microsoft, was hailed as a 4D Chess Champion/Galaxy Brain when it became public knowledge in 2022 that Microsoft had invested in OpenAI. Around the same time, Google realised that the rug was about to be pulled from under its feet and hurriedly released Bard. Luck wasn’t with Google and the Bard demos went awry. This led to more panic within Google and its stock.
Given the probabilistic nature of LLMs, it shouldn’t have been surprising, especially in the early days of LLMs. Back in 2022-2023 though, it seemed that Google was about to be disrupted as ChatGPT gained exponential traction. Nadella publicly stated that he wanted to make Google “dance” and it looked like he was winning.
Microsoft’s bravado and Google’s urgency was deceiving. Google had invested in AI for over 10 years and it had structural strengths in hardware and software that made it the natural leader of AI.
Google hardware such as its custom chip Tensor Processing Units (TPUs) were more efficient than Nvidia’s for workloads that Google needed (Search, Advertising) etc. Google also relied on rack scale infrastructure (lots of chips connected with lots of networking) that Nvidia adopted in Mar 2024 when it released NVL72. So in some ways, Google was actually ahead of Nvidia’s hardware given that it had been using AI in its workloads since 2016.
Google’s AI software division was also world class and the natural habitat of AI talent. Founders and key employees in all the major AI companies (OpenAI, Anthropic) all worked at Google at some point. The paper ‘Attention is all you need’ that kicked off the LLM revolution was written by Google employees who all later became AI luminaries.
Microsoft was always scrambling to catch up and the investment in OpenAI was a desperate attempt to get close to the starting line when the AI race would begin.
In mid-2026, the roles have reversed. Google has successfully staved off the ChatGPT challenge with Gemini and it still retains supremacy in Search. Microsoft has decided to step back from its relationship with OpenAI, scale down its funding of OpenAI’s hardware and support Open Source models.
Mark Zuckerberg (Meta) and Arvind Krishna (IBM) have also supported the adoption of Open Source models. Alex Karp of Palantir made colourful exhortations advocating Open Source models as a counter to handing over all the effing family jewels (data) to the big AI companies (Anthropic, OpenAI).
With Open Source models now closing the gap with the mighty Mythos and even the prescient Satya Nadella foreseeing an Open Source future, we come full circle back to the original question – can Frontier AI (Anthropic, OpenAI) companies ever be profitable for long and do we need a $ trillion hardware buildout?
The futility of predictions
The same old questions pop up repeatedly – Open Source vs Frontier. In a few months, we see the same old answers – Frontier Models doing wondrous things, Open Source catching up shortly.
Our human brains are bad at predictions, especially when unintuitive concepts like exponential growth or probability are at play. In this blog, we leave confident prognostications about how AI will or will not change the world to the content creators and the tech gurus.
We listen to Jensen Huang, who has been remarkably correct all these years (Scaling Laws, Hardware buildout, move from copper to Optics in networking etc). We listen to the AI doomsayers or the AI Nirvana bunch and nod along. We pay careful attention to informed criticism from the real experts and myth busters such as Gary Marcus. But we shall never allow ourselves the luxury of an opinion or heavens forbid, make a prediction.
Steve Eisman was one of the famous ‘Big Shorts’ of 2008. He, alongside Michael Burry and others, made a lot of money shorting the markets during the housing bubble. In one of the CNBC/Bloomberg interviews (not sure which), he exclaimed “Have some humility” when some typical market event didn’t lead to a much feared cataclysm. Here is a man who made his money and reputation being bearish but avoiding the trap of believing he can predict future bubbles.
Look East
Let’s try a different approach to the Frontier vs Open Source question.
Let’s imagine a country where Open Source models are as good as Frontier Models & foreign AI Models as well as Nvidia hardware are outright banned. Wouldn’t that settle the debate once and for all?
Such a country does exist – China.
Look East > Frontier Models (Claude, ChatGPT) in China
First, we look at Frontier Models in China. As we saw, Chinese Open Source models are for all practical purposes as good as the Frontier models. In Jul 2026, the new kid on the block Moonshot from China, released Kimi K3. This new model matched Mythos and even exceeded Mythos in certain aspects. (Frontend Code Arena benchmark).
Mythos was one of those models that was too good to be released to the wider public. The White House had even temporarily prohibited the use of Mythos by ‘foreign nationals’. And yet, Open Source from China caught up in a few months.
(The White House ban on foreign nationals accessing Mythos created an amusing situation where Andrej Karpathy, new Anthropic employee and the undisputed heavyweight champion of coding, couldn’t officially use the model.)
One would assume that if Chinese Open Source models are world class and practically free for enterprises, nobody would use the top Frontier Models in The Middle Kingdom. Especially if the US government could arbitrarily ban such models.
The surprising answer is that the Frontier Models are just as popular in China, even though they are technically banned. An elaborate ecosystem has emerged that makes these top foreign models available via API proxies or ‘Transfer Stations’ for as low as ‘10% of the official price’.
The blog ChinaTalk outlines how these Transfer Stations use ingenuity to outwit the smartest companies in the world and access banned tokens from OpenAI/Anthropic.
“Meanwhile, every layer of control frontier US AI companies have added (geoblocking, phone verification, credit card requirements, and now live biometric KYC checks) has produced a corresponding layer of evasion infrastructure“. Source
These Transfer Stations take extreme measures to work around the KYC requirements using Agents (the Human variety).
“Agents travel to lower-income countries in Africa or Latin America to recruit real individuals willing to complete in-person verification. The Worldcoin black market offered a documented precedent, with iris scans harvested from KYC merchants in Cambodia and Kenya, sold for under $30.” Source
Frontier Models (mainly from US/West) are thriving in China despite barriers. These barriers come from Chinese & US governments and US AI companies (to prevent distillation). Despite these restrictions, the demand for premium AI with premium pricing is increasing.
The situation in China has proven that premium AI models will continue to charge seemingly exorbitant amounts, even if access to these models is limited and cheaper alternatives are available. An entire economy is created, just to provide access to the top AI models.
Look East > Frontier Hardware (Nvidia/AMD) in China
The latest Hardware from Nvidia (GPUs) are also banned by US government for export to China and heavily ‘discouraged’ by the China government.
Both governments are playing a very delicate dance. The US Government wants to ensure that the western models remain supreme and have blocked the sale of the most advanced chips from Nvidia and AMD to China. An outright ban might hamper Chinese AI models in the short term but also risk the development of a Chinese chip ecosystem that might eventually match the best chips that US has to offer. Regulatory Uncertainty has been the US response to this dilemma – allow sale of some chips to China (never the best ones) but also keep them guessing by frequently banning the sale of the ‘nerfed’ chips as well.
The Chinese government also has conflicting short-term and long-term objectives. It wants to ensure that its AI capabilities are as good as US (get as many Nvidia chips) but wants to develop a local tech ecosystem in AI hardware that makes it independent of US (discourage Nvidia chips in favour of local chips). If the US government allows sales of some Nvidia chips to China at some points, the Chinese government imposes its own ban on Nvidia but occasionally relents.
In this delicate dance, Jensen Huang is being kicked around like a football by both governments. (someone give him a raise, he just isn’t getting paid enough).
Given all this competitive policy opaqueness by both governments, can we expect the demand for Nvidia chips to die down in China?
The opposite seems to be happening. Just like we saw Transfer Stations crop up to meet the demand for Frontier Models, we are witnessing clever workarounds to access the forbidden Nvidia chips.
During the early days of the ban, we saw Nvidia chips being smuggled as seafood, with a few crabs crawling about on top of the contraband.
In the current lucrative market for Nvidia chips in China, individuals have now come up with more sophisticated solutions to meet the demand. Chips are being sold to Singapore/Malaysia/Taiwan and then sent to the Chinese mainland. Shell companies have mushroomed that re-route shipments to China. New loopholes on the export ban are exploited as soon as the old ones are plugged. A $2.5 billion ‘heist‘ was discovered in Mar 2026 that shipped entire Supermicro servers containing Nvidia chips to China.
Anecdotes and hearsay? Agreed. So how do we assess the demand for a product where the supply is artificially tamped down?
We look at the Black Market. In the Black Market for Nvidia hardware, the prices are 1.6-3x the official prices.
- A100, practically ancient chips; 3x markup – These chips were first released 5 years ago. Ilya Sutskever, Geoffrey Hinton and other grandees of Artificial Intelligence were playing around with them while developing the first LLMs.
“The price of an A100 server has roughly tripled since late last year,…” Source - H100 chips, the workhorse of AI; 1.6x markup – These chips are one generation older than the current Blackwell chips. The value of these chips should depreciate 30% every year since their launch in Mar 2022.
Instead, “A single H100 that costs roughly $30,000 through legitimate channels commands $50,000 or more on the black market.” Source - B200 chips, shipped in racks, 1.5x markup – These are the latest Blackwell class chips.
“A rack containing eight B200 AI GPUs costs roughly $420,000 to $490,000 in China, about 50% above US prices.” Source
Note that Nvidia is famous for having Gross Margins of around 75%. This is a ridiculously high number for a company that is shipping metal boxes with wires. Even Software companies with near zero marginal cost (you don’t even need to make and courier CDs nowadays) can’t match these Gross Margins. In China, Nvidia products are being sold for 1.6x to 3x, on top of the 75% juicy margins that Nvidia charges.
In the movie Steve Jobs (2015), Woz tells Jobs “It’s not binary, you can be decent and gifted at the same time”. In podcasts nowadays, you will hear ‘Two things can be true at the same time’.
The Chinese AI market has shown us that it is possible for insane demand to exist for Claude when Open Source models exist that are practically free. A thriving black market for Nvidia also exists – despite efforts of the two most powerful governments in the world and the presence of cheaper domestic alternatives (Huawei Ascend).
Conclusion
We are living in an ExPro world (exponential and probable). In this brave new world, paradoxes abound. Prediction is futile.
As our world changes in anticipation of the AI dawn, the following will recur:
- AI progress, already at an exponential level, will pick up even more pace. AI models are at the threshold of Recursive Self Improvement – they will learn how to become smarter on their own.
- The general public will alternate between considering AI as Hype to Yawn to Doom.
- The AI buildout will always seem excessive. The smartest minds in the world will be accused of throwing money at AI like drunken sailors. They will continue to do so anyways because of the loaded gun at the end of the race.
- New Open Source models will keep emerging. They will threaten to render those fancy Frontier Models from Anthropic/OpenAI obsolete. The markets will panic and question the trillions spent on building data centers when the same can be accomplished using a few stacks of Apple Mac Mini.
- Every few months, the Frontier Models will hit back and do something so unexpected that the end of the world will seem near. The US Stock Markets will rally and everyone will be relieved at not being made to look stupid after investing trillions in AI.
- Nvidia has 90% share of the world’s most lucrative market share. By definition, its market share will decrease as the Galaxy Brains of the world try to get a piece of the action. Its chips will always seem expensive but Nvidia has perfected the art of pricing in the Gaming industry. The Gaming Industry is highly competitive and has Winner-takes-all characteristics. The AI industry is similar and Nvidia commands top dollar as it understands the competitive instincts and fears of the key players.
- We haven’t reached Wave 4, Robotics yet. Robots with intelligence – try predicting how that changes the world. Developments in Quantum Computing will open up avenues that even Agentic AI wouldn’t dare.
The only thing we can do is watch the Scaling Laws. If the AI models stop getting better even if they were trained on exponentially larger amounts of compute, the party is about to wind down.
Outtakes – This article took forever as newer developments kept contradicting or corroborating the delicate points being made. At the end, I had to follow the advice of the Duke of Wellington (of Waterloo and the many statues fame) – “Publish and be damned”. His response was to a blackmail attempt but applies for writers who tarry too long with their articles. Some developments of note and interesting tidbits below.
- The Brave New World by Aldous Huxley described a society that was ‘intelligence-based social hierarchy’. The Tokenmaxxing bros felt they had an year or two before AI would become pervasive and consign them to the ‘permanent underclass’.
- Circular financing got downright dizzying. Nvidia ended up owning a small part of most or probably all its competitors. Others also followed suit and now everyone owns a bit of everyone else. ‘Vendor financing’ made a comeback after 2000. Credit Default Swaps (CDS) instruments from 2008 crisis are being watched intently for any signs of defaults. Technological events of this magnitude have always led to a bubble of epic proportions. This one seems no different.
- Jonathan Ross, who explained Jevons paradox, has joined Nvidia.
- As I was writing this piece, a new Open Source model Kimi was released and the global markets panicked again. High flying bros running Tech Funds had to liquidate their holdings to wily old foxes who didn’t understand AI but understood markets.
- The Open Source/Weight models such as Kimi have become so big that they will also need Nvidia hardware to run.
- Newer Frontier AI models are being delayed. Official line is that they are too dangerous. AI sceptics possibly hint at any lack of progress and a break in Scaling Laws.
Sources
- Tokens
- Fall in Token cost – Blackwell 35x lower than Hopper Nvidia blog
- Token Explosion – Azeem Azhar
- Anthropic code per developer – “The amount of code contributed per developer is eight times higher in the current quarter than the average to the end of 2024.” Azeem Azhar
- Exponential
- RSI – Recursive Self Improvement – AI now improving itself
- Mythos
- Hype – Gary Marcus
- Frontier
- Open Source
- The Substitution Wave in AI – Tomasz
- ‘Six Chinese open source models now sit at or near frontier quality at 5x lower cost’ Podcast Alpha
- “A fascinating deep dive into the Chinese black market for Claude tokens, which trade for up to 90% off list price.” Original S; S
- Open Source, Open Weight or neither – Culpium
- Jevons Paradox and Cost
- Chip Cost – Fall in prices “Noyce cut the price of the IC from $120 to $15, an 87.5% drop, while still charging NASA and MIT premium prices to fund scale. “Noyce slashed prices, too, gambling that this would drastically expand the civilian market for chips, Chris Miller wrote in Chip War. “In the mid-1960s, Fairchild chips that previously sold for $20 were cut to $2. At times Fairchild even sold products below manufacturing cost, hoping to convince more customers to try them.”” S
- AI Revenue
- OpenAI – Azeem Azhar