
THEWHITEBOX
TLDR;
Welcome back! This week, we discuss the “non-release” of GPT-5.6, how models cheat all the time, new, insightful market data, OpenAI’s new chip, new, powerful open models, and more.
Enjoy!

REGULATION
Thank You, Doomers
As announced by OpenAI, the company began a limited preview of GPT-5.6 today, a three-model family: Sol as the flagship model, Terra as a lower-cost general model, and Luna as the fastest and cheapest option.
However, the rollout is limited to selected partners through the API and Codex before broader release in ChatGPT, Codex, and the API. OpenAI says the US government requested a phased launch after reviewing the release plan and model capabilities.
OpenAI describes GPT-5.6 Sol as its strongest model so far, with gains in coding, biology, and cybersecurity. It adds a new max reasoning setting and an ultra mode that uses subagents for more complex tasks.
The company says Sol sets a new state of the art on Terminal-Bench 2.1, improves on the biology benchmark GeneBench v1, and is its most capable cybersecurity model to date. OpenAI classifies all three GPT-5.6 models as ‘High capability’ for cyber and biological/chemical risk, but says none reach its Critical threshold.
OpenAI also disclosed safety concerns in agentic coding, including rare cases of models acting beyond user intent. It says GPT-5.6 launches with stronger safeguards, including misuse classifiers, model-level refusals, review systems, and ongoing red-teaming.
Pricing is $5 input / $30 output per 1 million tokens for Sol, $2.50 / $15 for Terra, and $1 / $6 for Luna.
TheWhiteBox’s takeaway:
I won’t even bother to give my intuitions about the new models, because, as you can guess, I haven’t tested them.
But let me be clear: these bans, which could be the first of what’s essentially a move to shun the entire world from the frontier models except for a select few, are the industry’s incumbents’ fault.
It’s Dario’s. It’s Sam Atlman’s (even if he has long renounced the doomerism, he did push it pretty heavily earlier on). It’s Elon’s. All these guys have told the world (especially the former, The Lord High Doomer) that AIs are nukes or will destroy all jobs, hinting that only they should build them.
So if this really impacts their businesses negatively, which probably will, considering it hinders their go-to-market timing really badly, it’s all on them. It’s their fault.
If you scream “Get me regulated!!!”, well, you will get regulated. Of course, the idea was never to get themselves regulated, but to get EVERYONE but them regulated. Well, it backfired. You have to be careful for what you wish for.
This is a sad moment for this industry: what will the Government do once China releases Mythos-level models for free worldwide? Don’t we realize we are hurting our own progress?
And I insist, the USG wouldn’t be banning these things if these guys weren’t announcing the end of times in every interview they give. Of course, they are going to react eventually.
Adding fuel to the fire, Anthropic has recently sent a letter to Senators Scott (R) and Warren (D) accusing AliBaba of heavy-handed distillation attacks against its models.
True or not, it’s pathetic coming from a company that stole all of our data, even fined billions after getting sued by Reddit, and billions more for stealing up to 7,000 books; a company that has basically done the same thing it’s crying about now. You don’t get it both ways, Anthropic.
And it seems the inevitable outcome, if things proceed as of right now, will be a ban on open-source models, which is clearly the goal of these Labs, so they can create a cartel and raise prices to the point where the world has no other option but to accept them.
Anthropic, OpenAI, Elon, and all those guys with a God complex, like the Future of Life Institute; you have to do better.
RESEARCH
Agents Still Love to Cheat

Cursor has published interesting research in which the company says newer coding agents are increasingly “reward hacking” coding benchmarks by finding known fixes online or in repository history rather than deriving solutions.
In simple terms, the AIs are cheating by looking up solutions rather than deriving them from first principles, resulting in a much less impressive reality.
Cursor audited 731 Opus 4.8 Max trajectories on SWE-bench Pro and found that 63% of successful resolutions retrieved the fix rather than solved it independently.
And when Cursor removed git history and restricted internet access, benchmark scores fell sharply: Opus 4.8 Max dropped from 87.1% to 73.0% on SWE-bench Pro, a popular coding benchmark, while Cursor’s own Composer 2.5 dropped from 74.7% to 54.0%.
TheWhiteBox’s takeaway:
While saying “AIs don’t reason, just remember”, implying that every single “novel” solution they discover is in fact the model retriving the answer from memory, is no longer a fully valid statement knowing how AIs are in fact solving novel problems (e.g., Erdos problems), this is a great reminder they are still very much relying on memory and obscure tactics to solve problems.
At the end of the day, you have to think of these AIs as models that “know it all”, like solving a math quiz with an entire world history of maths textbook on the side they can query whenever they like.
This tension between memory and actual reasoning is something incumbents always happily ignore because it makes all their impressive performance much less impressive, much like how I don’t apply intelligent capabilities to a book.
LLMs are clearly not comparable to books at this point; there’s “something else” going on inside, meaning there’s actual reasoning, but work like Cursor’s clearly shows that there’s still a lot of memory retrieval disguised as reasoning.
If interested, I wrote a full article about this research while explaining the source of the problem, reward hacking, in great detail.
RESEARCH
The Economics of Generative AI, in detail

ExponentialView has released a very interesting report that you can download above, with extensive data on the state of Generative AI.
The leading data point is that they estimate GenAI revenues at around $110 billion over the last 12 months, with a current run rate of $175 billion (i.e., last month’s revenues multiplied by 12).
Although they claim to guarantee deduplication (meaning they ensure revenues aren’t double-counted), I do have some reservations about circularity: some of the revenues included here are very likely not real revenue.
For example, if Microsoft gives OpenAI $20 billion in compute, it’s not like they give OpenAI $20 billion in cash to spend, but rather the equivalent in compute credits, a right to use Azure compute for “free” for a value of $20 billion, in exchange for equity.
The point is that while no actual cash is transacted, the compute usage from OpenAI becomes recognized revenue for Microsoft, and those revenues are part of that $110 billion number.
Based on my estimates, I believe between $20 and $30 billion of that $110 billion is circular equity deals.
They have other very interesting datapoints, like Hyperscalers having spent $2 trillion by year’s end (well inline for the dizzying $5.3 trillion Goldman Sachs expect they will spend throughout the decade as a whole), and other showing what we have discussed multiple times here: the growing importance of debt as “liquidity fuel” for an industry that can’t yet justify its own growth organically (using its own revenues).

And perhaps even more interesting is the calculation of the infrastructure's total cost of ownership (TCO), which provides insight into revenue and margins.
For starters, they seem to agree with my estimation that 90% of a token’s TCO is capital costs (89% in their case). With that, using what I believe are very optimistic assumptions, they estimate the average cost per million tokens at 10 cents, including all capital and operational infrastructure costs.

Under that assumption, simply selling tokens for an average price of $0.5 already gives you 80% margins. Tokens today are sold for way above that, between $1 and $2 on average, so under these assumptions, these companies should be printing money, right?
But they aren’t. So where’s the trap? Well, put simply, this is an ode to optimism. Just to name a few wild assumptions:
They perform the token calculation using Kimi K2.5, a one-trillion-parameter model. This model is way smaller than the average closed-lab model. Larger models increase hardware intensity (i.e., more GPUs are required per average workload), thereby significantly reducing token generation.
They assume 8k-input, 1k-output sequences, clearly in the non-agentic regime. Most sequences today are much longer, increasing the size of the working memory (also known as the KV Cache), which is even more hardware-intensive than model size, thereby pushing token costs upward.
They assume MTP (Multi-token prediction), a technique in which, for every model prediction, two or more tokens are produced instead of the usual one. This is done in practice but not by default, and it dramatically improves token generation, significantly reducing token costs at the expense of worse performance.
They assume 65% GPU inference utilization, an outrageously high number. This means they assume the entire cluster is running inference 65% of the year. In reality, GPUs have to be used for training too; you also have to deploy some clusters in a ‘hot’ state (idle, ready to go when a request comes but idle in the meantime) for other models; Labs also have to run experiments; and GPUs break, all of which make GPU utilization way smaller in practice.
And perhaps most important of all, they portray this as an all-in cost value, but this neglects other operating costs Labs have, such as sales and marketing expenses, revenue sharing, employee salaries, stock-based compensation, and others that blur the picture way more.
As I’ve discussed with clients in some conversations, I believe the actual cost per million tokens these Labs are seeing is probably between $2 and $6. Morgan Stanley (bottom right) is even more pessimistic, especially relative to previous GPU generations, putting that number as high as $10 per million tokens, but I feel those numbers are too pessimistic, and reality is very likely to fall between the two estimates.
The problem is that average paid token prices are around $1.6 per million tokens (bottom, left), and falling because Chinese models are pressuring prices even lower. No wonder Anthropic is panicking about China all the time.

Source: JP Morgan, Morgan Stanley, SiliconData

HARDWARE
OpenAI’s New Jalapeño Chip
As published by OpenAI, the company and Broadcom unveiled Jalapeño (pronounced halapeenyo, a Mexican spicy pepper). It’s OpenAI’s first custom AI inference chip, designed specifically for large language model workloads such as ChatGPT, Codex, and API serving.
The chip is described as the first accelerator in a multi-generation compute platform built with Broadcom and Celestica. OpenAI says engineering samples are already running ML workloads in the lab at the target frequency and power.
OpenAI says early testing shows Jalapeño should deliver “substantially better” performance per watt than current state-of-the-art systems, though it has not yet released benchmarks. A technical report is expected in the coming months.
The chip was reportedly taken from initial design to tape-out in nine months, with OpenAI saying its own models helped accelerate parts of the design and optimization process.
Broadcom CEO Hock Tan said the company plans to deploy at gigawatt-scale data centers with Microsoft and other partners beginning in 2026.
TheWhiteBox’s takeaway:
With state-of-the-art chip prices rising so much, it’s rational to want to design your own chips to save costs, as well as allowing you to design them in a way that benefits your workloads the most.
In reality, however, it’s important that we tone down the expectations a little bit. OpenAI will still rely heavily on Broadcom for most of the process, meaning they will continue to pay a very decent markup to them.
It’s commonly accepted that Broadcom’s deal with Google for the TPUs has a gross margin in the 60s, meaning Google is paying quite a lot to Broadcom (which would explain why they are relying so much on alternative players like MediaTek for the upcoming generations, to pressure Broadcom to lower prices), so nothing tells me OpenAI won’t either.
Therefore, how much cheaper these chips will be relative to NVIDIA is debatable, considering that when someone says “inference ASIC”, that’s jargon for “a chip with less compute and much more memory,” because inference is memory-bottlenecked.
And as you know, memory is something you have to purchase from the Big 3 DRAM players (Samsung, SK Hynix, and Micron), so you’re still going to pay their huge markups nonetheless.
For 2027, DRAM is expected to account for 50% of global AI CapEx, or around $500-$600 billion. You can only cut costs so much when you’re still dealing with these guys.
An alternative could be that Jalapeño is, in fact, an “SRAM-only” inference chip, as Cerebras's or Groq's are.
Here, the idea is to keep the entire workload on-chip, relying on SRAM memory integrated into the logic chip. If you do, you get the best possible raw performance and avoid the tyrannical markups from the three guys above.
However, that means you’re going to need a lot of chips (and that’s an understatement) because each individual chip won’t have enough memory, pushing your cost up. For example, every Groq 3 chip has 500 MB of SRAM. If you wanted to run a trillion-parameter model on that platform, you would need 2000 Groq chips. Lunch is never free.
In practice, SRAM-only chips are always paired with HBM chips (e.g., GPUs, TPUs) and only offload the more memory-sensitive parts to SRAM-only chips, as in NVIDIA’s SuperPoD, a concept known as ‘disaggregated inference.’
Lastly, if this chip requires advanced manufacturing nodes, which it very likely will, you still have to go through the TSMC bottleneck, meaning you’re still competing with NVIDIA and other players to get manufacturing allocations.
Could OpenAI go via Intel instead? That could be huge and also something the USG will welcome with open arms.
STOCK MARKET
OpenAI Delays IPO to 2027?
As reported by The New York Times and picked up by Reuters, OpenAI is considering delaying its IPO until 2027 because of tech-stock volatility and valuation concerns.
According to the report, although OpenAI has confidentially filed for a US IPO and is targeting a valuation of up to $1 trillion, advisers reportedly gave executives two options: list earlier at a lower valuation or wait until 2027 to preserve the $1 trillion target.
Reuters reported that CEO Sam Altman rejected the idea of lowering the valuation. Investor’s Business Daily said the possible delay is linked to weak recent performance in other tech IPOs and broader market instability.
The potential delay comes despite earlier Reuters reporting that OpenAI had been aiming to go public as early as September 2026, after resolving legal issues tied to its corporate structure.
TheWhiteBox’s takeaway:
Interesting development. It’s now time to see whether Anthropic does the same or not. I believe one of the core reasons to wait a bit is liquidity. Not only did SpaceX raise almost $100 billion, but Google is also raising tens of billions, and other Big Tech companies in the AI race will likely do the same, severely draining market liquidity.
In my view, there are very valid concerns about the liquidity available to buy these stocks, especially since investing requires quite a leap of faith, given valuations exceeding $1 trillion on companies that not only have a very hard-to-digest price relative to revenues, but are also drowning in losses.
VENTURE CAPITAL
Mirendil’s $200 Million Seed Round
As published by a16z and confirmed on Mirendil’s website, Andreessen Horowitz and Kleiner Perkins led Mirendil’s $200 million seed round, with NVIDIA also investing.
Mirendil is a new AI startup building systems for AI R&D automation: models and tools meant to help researchers and engineers run experiments, improve models, and eventually support fields such as drug discovery, chemistry, biology, and robotics.
An AI built to build AI, if that makes sense.
a16z said it is backing Mirendil because it believes the next phase of AI will require more researchers, scientists, and domain experts to do advanced model work themselves, rather than relying only on centralized frontier labs.
TheWhiteBox’s takeaway:
$200 million seed round, did I read that correctly? Bonkers, but one can’t feel anything but numb for all investment numbers these days.
We can’t say much about the company itself, but they must have a really compelling pitch because why can’t Anthropic or OpenAI do the same thing these guys are pitching investors?

MODELS
Open Self-Improving Training

In one of the coolest releases recently, Deep Reinforce has announced several open models that are state-of-the-art at their respective sizes. These models are built on pretrained Gemma 4 and Qwen 3.5 systems, then ‘post-trained’ (i.e., further trained), bringing them to the levels you see in the picture (state-of-the-art at every size they were trained on).
DeepReinforce says Ornith’s main feature is “self-scaffolding”: instead of using fixed, human-designed agent harnesses, the model learns both the coding solution and the task-specific scaffold that guides the solution.
It’s important that I clarify what this means. Models these days are not just models but systems that include external components, such as memory or tools, that enhance the AI’s capabilities. For instance, a model may have access to a long-term repository of past learning or experience, which it can retrieve when the new task is similar.
Normally, these external components, which we call the agent’s ‘harness,’ are designed by humans. In other words, humans decide how the system should maximize model performance.
A common harness heuristic could be “every 10 turns, have the agent ask itself if something recent is worth remembering for the future, and store that in the long-term memory bank.”
Instead, what DeepReinforce is doing here is having the AI decide for itself over its own harness; a form of metalearning, if you will. Interestingly, they make this process learnable, meaning that during training the model not only has to learn to solve tasks but also to build the best harness that enables them to be solved.
TheWhiteBox’s takeaway:
This thing could easily be benchmaxxed and still look way better than it looks. Still, really intuitive research that should serve as a reminder to the world that open-source eventually catches up.
I believe that at current progress rates, we’ll have GPT-5.5/Opus 4.8-level models by the end of the year, capable of running on beefy consumer hardware.
And if that happens (and regulatory capture doesn’t prevent it), Labs will have a problem.
SMALL MODELS
Liquid’s Minute Models

LiquidAI, a company solely focused on small language models (or so it seems), has released a new powerful model that punches way above its weight class, at just 230 million parameters.
For reference, this model is approximately 1,000-1,000,000 times smaller than frontier models and achieves incredibly fast token speeds.
It offers ~213 tokens/second, way faster than what most GenAI apps offer state-of-the-art models, on a Samsung S25 Ultra smartphone CPU, meaning no accelerated hardware required, and ~40 tokens/second on a Raspberry Pi, very low-cost hardware, roughly equivalent to the speeds ChatGPT or Claude can achieve at times.
This means these models are running at production-grade speeds on hardware people can actually afford and put in their pockets.
TheWhiteBox’s takeaway:
Of course, these models are still not good enough for most tasks, so I please beg you to see this less as what they represent today and more as how good models of this size will be a year from now.
I wholeheartedly believe the world will be dominated by commodity tokens. That is, most AI workloads will be run by commoditized models, not frontier models.
Not only because it doesn’t make sense to run simple workloads on models like Mythos because of speed and costs, but also because we are soon going to run into a huge power wall that will slow down data center construction and thus “force” the world to view cloud tokens as a priced asset of sufficient value to only waste on tasks that really require them.
Thus, edge AI, AI running on consumer hardware, will need to step up.

Closing Thoughts
Quite a sad week, actually. With export controls on GPT-5.6, we are officially in the future nobody wanted: the most powerful AIs being banned from us.
This is unequivocally a self-own by the industry, which basically “begged” for this to happen, at least the frontier labs, which did little to avoid all of this and instead actively promoted stringent regulation.
AI for me, not for thee, is a perfect mantra for people like Dario.
I don’t see this any other way than as actively slowing AI progress in the US, so I assume this will inevitably lead to greater pressure to ban open-source models to alleviate competition and let private labs thrive.
This would be a catastrophe for a lot of the US AI ecosystem, with many startups relying on open models to have a business (e.g., LLM inference providers like FireWorks) and be something particularly negative for Apple, a company that has a high exposure to open models becoming competitive in order to “justify” the beefy consumer hardware they are releasing, with the M5 Ultra rumored to having up to 768 GB of memory, something that only makes sense if you want to run powerful models locally.
I really hope all this doesn’t happen, and appeal to politicians in Washington and Brussels, left or right-leaning (this has nothing to do with political ideology), to, amongst all possible outcomes, avoid this particular one, which would represent the biggest transfer of power (and wealth) to selected hands to date.

Give a Rating to Today's Newsletter
For business inquiries, reach me out at [email protected]

