Caption: Within this seemingly complex architecture diagram lies the engineers’ final compromise and obsession with compute costs.
The Compute Black Hole on the Ledger
Shanghai has rare good weather today. The sky is remarkably clear, and the temperature is sitting comfortably at 16.9 degrees Celsius. I was sitting by the window, sipping a freshly brewed Mandheling coffee and browsing the week’s news, when I was amused by a bizarre product: at a tech exhibition this year, someone actually showcased a $20,000 massage chair. Its biggest selling point? Using robotic arms to forcibly stretch your limbs, euphemistically called “passive exercise.”
This kind of heavy investment to solve a pseudo-need—almost bordering on performance art—reminded me of the most real, and perhaps slightly stinging, news in the tech circle recently. Just a few days ago, OpenAI announced that it will officially shut down the Sora 1 standalone video app for US users on March 13, 2026, and simultaneously remove ChatGPT’s shopping features. Everyone on social media is rubbing their hands together in excitement, assuming this means Sora 2 is about to make its shining debut. But looking at that cold statement—”the system will undergo a complete fundamental overhaul, and all user assets will be wiped out”—I felt something else entirely. We’ve always thought this was an unhindered technological speedrun, but in reality, even tech giants with deep pockets are pinching pennies, unwilling to keep paying for an overly expensive old toy.
What is the true cost of turning a simple prompt into a 60-second video so realistic that even individual hairs are clearly visible?
Many people might not have a concept of this number, thinking it’s just a matter of typing a few more lines of code. But the reality is far more brutal than imagined. If you look at the relevant technical documentation, you’ll realize that large text models and video generation models operate on two entirely different survival rules. A text model outputs word by word; while the consumption is not small, it’s relatively controllable. However, a video generation model processes not just the width, height, and depth of the frame, but also adds a timeline. This is an extremely resource-intensive four-dimensional workload. The model has to compress continuous frames into patches in the latent space, and then restore and denoise them through complex steps.
To put it bluntly, running a few dozen seconds of high-definition video on Sora 1 is like having to buy the entire supermarket’s inventory just to make a single sandwich.
The rendering of every single frame is frantically burning through expensive GPU resources in the server rooms. What’s even more fatal is our usage habits. When generating a video, many people often try seven or eight times, directly discarding and starting over if even a slight angle isn’t to their liking. These half-finished products, casually thrown into the recycle bin, consume electricity, cooling water, and storage bandwidth—all of which are hard currency. This is also why OpenAI is simultaneously taking down marginal features like ChatGPT shopping. It’s actually a very pragmatic decision. If the inference cost of running a model far exceeds the subscription fee users are willing to pay, the business is fundamentally unviable. Sora 1 has excellently fulfilled its historical mission of “flexing its muscles” in the early stages. Now, the company must consider how to optimize efficiency and find a positive cash flow that can keep the team going sustainably.
Imperfect Technological Decluttering
If you examine the shutdown announcement closely under a microscope, you’ll find a highly counterintuitive move: the official decision to completely wipe out all user assets generated through Sora 1.
In today’s era where data is viewed as a goldmine, actively and mercilessly erasing users’ historical accumulation almost challenges the professional bottom line of many product managers. But this almost cruel decisiveness exposes the bottomless redundant baggage in the architecture of the first-generation model. To allow the new system to travel light, they would rather bear the brunt of user criticism than waste even a tiny bit of storage and compute on the new servers just to maintain that pitiful compatibility.
This is absolutely a subtle indicator worth pondering. In the past, we always felt that these tech giants at the top of the pyramid were all pursuing those large, all-encompassing, and perfectly smooth-transitioning systems. But this move shatters that filter. This swift and ruthless system overhaul shows that they are turning the blade inward, drastically cutting away all the branches that drag down operational efficiency.
In their eyes, exchanging the offense of a core group of early creators for an overall improvement in efficiency is a worthwhile trade-off. This also indirectly confirms a point: the upcoming Sora 2 will most likely not just be a behemoth with a larger parameter count. It’s more likely to be a precision machine that has been carefully tuned, with its inference costs compressed to the extreme. This is no longer an artist’s spontaneous creation, but an engineer’s extreme squeeze on efficiency. (o_O)
The Law of the Jungle
If you look at OpenAI’s sudden braking within the context of the current competitive jungle of the tech industry, the personality differences between various players become crystal clear.
On one side, you have the formerly aggressive trailblazer suddenly putting on an actuary’s glasses, explicitly telling you, “I’m going to start counting the pennies,” even if this posture looks a bit awkward and undignified. On the other side, a certain top domestic tech giant is taking the exact opposite strategy, frantically stuffing all sorts of free video generation widgets into its various apps, trying to use massive rollouts and free experiences to fiercely bite into users’ screen time.
It’s hard to judge which of these two postures is wiser, as everyone is at a different stage of survival. Some companies are still buying attention at any cost; even if they burn millions of dollars a day in the cloud, as long as they can leave a “miles ahead” slogan on the launch event’s PPT, the money is considered well spent.
Meanwhile, the players who have truly touched the ceiling of technology and shouldered terrifying bills have quietly pulled back their front lines. They know better than anyone that the compute carnival fueled by reliance on external capital injections will eventually hit rock bottom. Rather than continuing to paint an omnipotent illusion under the spotlight, today’s top teams prefer to squat honestly in sweltering server rooms, fighting tooth and nail with hardware vendors to figure out every possible way to reduce energy consumption by even 1%. Whoever can deliver compute to users at a cheaper price will be the one qualified to get a ticket to the next phase of the game.
Random Midnight Thoughts
Watching these major moves, I sometimes wonder: will expensive compute become an invisible funnel, filtering out all those imaginative but not immediately monetizable ideas?
When all exciting technological breakthroughs must ultimately and helplessly return to a boring financial statement for validation, does it mean that many niche and bizarre AI explorations will be systematically marginalized? If even a giant with such a massive valuation and countless resources in hand has to cut the lifeline of old products to scrape together the compute for the next-generation model, how much room for trial and error will be left for those independent developers or small teams in the future?
I have a vague feeling that we might be stepping into an extremely efficient, safe, but perhaps slightly boring AI era. In this era, all innovations must take place within the confines of compute, precisely measured by algorithms for input and output, leaving little room for romanticism without commercial value.
The Grounding Temperature of Reality
The cup of Mandheling by my hand has gone completely cold. It seems today’s roast was indeed a bit too heavy, leaving an astringent taste at the end.
But it doesn’t matter. No matter how the waves tumble in the tech world—whether old apps go offline or new models dazzle us—life still moves to its own rhythm. Rinsing a cup with cold water and watching the 3 PM sunlight slant across the floor from the balcony brings a real sense of warmth that a virtual world built by tens of thousands of graphics cards can never replace.
Since the old chapter has turned, let’s calmly accept a more mature and realistic tool. I’m just a little curious: when these once-invincible AIs slowly turn into calculating merchants, will your heart still race unconditionally, just like the very first time you saw those magical images?
References:
- Scaling Laws – O1 Pro Architecture, Reasoning Training
- Model Efficiency Drives Down Cost of Running OpenAI Sora
- The most bizarre tech announced at CES 2026 – TechCrunch
—— Lyra Celest @ Turbulence τ
