In March, Shanghai’s sky is dotted with a few drifting clouds, and the 12-13 degree (Celsius) temperature is still a bit chilly. Just as I was about to order a hot Americano, I saw that Google officially released Stitch, a tool that’s been rumored for a long time.
Anyone who has been following AI design should know that this isn’t exactly a brand-new project. Its predecessor is Galileo AI, which Google acquired and showcased at last year’s I/O conference. But this time, they’ve done a top-to-bottom refactoring, directly bridging the gap between natural language and UI prototyping.
They Call It “Vibe Design”
When I first saw the term “vibe design” coined by official sources, I was honestly a bit confused.
In traditional software engineering, design is an extremely precise discipline. I previously chatted with an interaction designer friend from a big tech company, and he complained that the most painful part of his day wasn’t a lack of inspiration, but rather spending 70% of his time on manual labor—building component libraries, calculating spacing, and stringing together a bunch of artboards into clickable interactive prototypes.
A major selling point of Stitch this time is tossing out these processes entirely. You no longer need to start by drawing wireframes; instead, you directly describe a business concept in natural language, or the kind of feeling you want users to experience. In the background, it uses the Gemini large language model to instantly translate these highly abstract words into high-fidelity UIs, smoothly stitching together multiple screens into a seamless interactive journey along the way.
Put simply, it’s like you used to go to the building materials market to pick out bricks and compare sizes yourself, but now you’ve hired an all-inclusive contractor. You just need to tell him, “I want a living room that feels safe,” and he handles the rest. This top-down workflow is actually a complete betrayal of the design logic centered around pixels and layers from the past decade. It attempts to make you surrender absolute control over details, shifting your focus to controlling “intent.”
Designers Shouting at Their Screens
Taking a closer look at their demo, there’s one detail that is particularly interesting.
This update includes a feature called “hands-free voice interactions,” meaning you can go completely hands-free, using your voice to command the underlying agent to tweak layouts or try out new versions. This does sound quite sci-fi.
But imagine putting this into a real office setting. In most companies’ open-plan cubicles, can you picture a row of designers wearing headsets, taking turns yelling at their screens, “Move that primary button a bit to the left,” or “Change this background to a warmer color”? (Rubs temples)
To be honest, spatial positioning and visual micro-adjustments are inherently difficult to describe accurately using linear language. Sometimes, something that takes half a second to fix with a mouse drag would require you to speak it out loud, wait for the AI to understand your intent, render it, and return the result. From an efficiency standpoint, this makes zero sense.
For a company as shrewd as Google, pushing such heavy voice interaction in a design tool surely isn’t just about saving designers a few mouse clicks. A more plausible explanation is that they are using this highly demanding multimodal understanding scenario to feed and polish the underlying LLM’s ability to process complex “visual-voice mappings.” Every voice command you make on the canvas, followed by your manual corrections, is perfect alignment data for them. This also explains why the official terms explicitly require users to be over 18, as this involves the massive collection of complex raw voice data and privacy authorizations.
To Put It Bluntly
If you stack Stitch up against the current pile of tools for a horizontal comparison, some issues become impossible to hide.
There’s no shortage of tools on the market that can generate interfaces from prompts. For example, there’s a slew of helper plugins within the Figma ecosystem, and then there’s Vercel’s v0.dev, which is incredibly popular in developer circles.
The core logic of v0 is very straightforward—it’s built for programmers. You input your requirements, and it spits out styled React code directly. You can copy, paste, tweak it a bit, and have it running in a production environment. Its interface might not be as flashy, but its engineering feasibility is extremely high.
In contrast, Stitch’s current positioning is a bit conflicted. It heavily emphasizes visual “high fidelity,” boasting about its ability to instantly link a bunch of screens into interactive prototypes, and even generating a QR code on the fly to send to your boss.
But there’s a catch you need to understand: UI design never ends after drawing the screens and presenting them to the boss. I previously saw survey data on frontend engineering indicating that over half of technical teams experience severe friction costs during the “design-to-code” handoff. If Stitch only generates a seemingly beautiful, flawlessly smooth “empty shell” prototype without exposing enough flexible architecture customization capabilities underneath, frontend engineers will absolutely break down when they receive these requirements.
To put it bluntly, a tool that dramatically lowers the barrier for upstream product managers to output designs while completely ignoring the survival of the downstream workflow can easily turn into a “sandwich” tool—one that leaves bosses thrilled during demos while developers curse in the background. A certain tech giant in China previously built a similar internal platform for one-click generation of high-fidelity interaction chains. It ultimately faded into obscurity after being collectively boycotted by frontend teams because the translated code was too messy and unmaintainable.
Imagination Constrained by Language
As a side note, regarding this entirely “language-driven” design paradigm, I sometimes wonder: could it actually be a bad thing?
Although human language is rich, it is inherently logical and structured. It’s incredibly difficult to use language to describe a visual form we’ve never seen before, one that completely breaks free from conventional understanding. Classic interaction designs of the past, like the slide-to-unlock on the original iPhone, or the subtle rubber-band bounce effect when pulling to refresh, were discovered through countless trial-and-error attempts with fingers on screens, relying on intuition and muscle memory. It’s hard to type into a black box, “Give me a sliding interaction as stress-relieving as popping bubble wrap.” AI cannot comprehend this highly emotional physical projection.
When we turn the starting point of all design into “describing a business concept” or “give me an interface like XYZ,” large models, in order to satisfy these explicit semantic requests, will always provide the statistically safest and most mass-market solutions. This is because its aesthetic is essentially the “greatest common denominator” calculated from tens of millions of design files on the internet.
Over time, all apps might start looking more and more alike, and the aesthetic of the entire digital world may shift towards an extremely standardized mediocrity.
Or maybe I’m overthinking it. Perhaps the future really will see the rise of highly literate “poet-designers” who no longer study color psychology, but instead spend their days figuring out how to use obscure yet precise metaphors to force large models into generating entirely new visual styles.
Actually, what I’m most curious about right now is how Figma’s PR team will respond tomorrow with an article facing this behemoth that attempts to swallow sketches, high-fidelity designs, and prototype interactions all in one bite. ¯\(ツ)/¯
References:
- Introducing “vibe design” with Stitch
- Exclusive: Early look at upcoming vibe design tool from Google
—— Lyra Celest @ Turbulence τ.
