AI, ML, and networking — applied and examined.
Can 1,000 AI Agents Predict the Future? A Deep Dive into MiroFish
Can 1,000 AI Agents Predict the Future? A Deep Dive into MiroFish

Can 1,000 AI Agents Predict the Future? A Deep Dive into MiroFish

MiroFish Interface — Real-time screenshot of multiple agents interacting in a simulated environment
This image roughly shows what it’s doing: a bunch of AI characters discussing the impact of public trust and emotions on group decision-making. To be honest, I was a bit stunned when I first saw this interface.

March in Shanghai, 6 degrees Celsius, cloudy—going out in just a hoodie will still make you regret it. Today happens to be International Open Data Day. Although it has no direct relation to what I’m about to discuss, the words “open data” reminded me of a recently active project on GitHub.

Built in Ten Days, Secured a 30 Million Investment

Let’s start with a rather magical fact: The author of MiroFish is a college student who built it in 10 days using Vibe Coding.

I previously saw his own retrospective post on LinuxDo, where he spoke very candidly—a college senior who finished his graduation project and wanted an internship, so he spent the last 10 days of summer vacation building a predecessor project called BettaFish (a public opinion analysis tool) using AI-assisted programming. As a result, it gained 20k stars within a week of being open-sourced, and his inbox was flooded with various offers and investment intents. Later, he joined Shanda Group and spent another 10 days upgrading BettaFish’s “analyzing the past” into MiroFish’s “predicting the future.” After watching the demo video, Chen Tianqiao decided to inject a 30 million RMB investment within 24 hours.

This story alone is enough for an article, but today I want to talk more about: What exactly is MiroFish doing technically? How credible is its claim to “predict everything”?

Doing an Old Thing in a New Way

If you strip away all the fancy packaging, the core logic of MiroFish is actually not complicated: give the AI a bunch of “seed information” (news, policy drafts, novel texts, financial data… whatever), let the system automatically extract entity relationships from it, build a knowledge graph, and then generate a large group of AI agents with independent personas and memories. Toss them into a simulated environment to interact freely. You stand by and watch what regularities emerge from their group behavior, and finally, the system hands you a prediction report.

Frankly speaking, this is Agent-Based Modeling (ABM)—an old method used for decades in social sciences, economics, and epidemiology. It’s just that the agents in traditional ABM are based on hand-written rules with rather mechanical behavior patterns, whereas MiroFish uses Large Language Models (LLMs) to drive each agent, allowing them to communicate in natural language and make decisions closer to those of humans.

Architecture diagram of the OASIS framework
MiroFish’s underlying layer uses the OASIS framework open-sourced by the CAMEL-AI team. This diagram shows its information propagation and agent interaction architecture. It’s capable of simulating millions of agents, although MiroFish’s current practical comfort zone is around 500 to 2,000.

There’s a notable detail in its tech stack: its simulation engine wasn’t written from scratch but stands directly on the shoulders of the OASIS framework open-sourced by the CAMEL-AI team. OASIS itself is a very interesting project; its paper was published on arXiv, claiming to simulate the behavior of millions of agents on social platforms like Twitter and Reddit—things like information spread, group polarization, and herd mentality. MiroFish essentially wrapped OASIS with a more user-friendly interactive interface and prediction report generation.

A Logical Hidden Concern

I chatted with a friend about this project, and he asked a very sharp question: Are the “group behaviors” simulated by these AI agents a true reflection of human society, or just an amplification of the biases inherent in Large Language Models?

This is actually a problem faced by the entire field of LLM social simulation. In February 2026, the Stanford team behind Generative Agents (the authors of the classic 2023 paper “25 AI Agents Living in a Virtual Town”) just founded a company called Simile, raising a $100 million Series A. Their data shows that AI-simulated questionnaire responses can achieve an 85% accuracy rate in replicating human self-reporting. Sounds good, but conversely, there’s still a 15% deviation—in scenarios requiring extremely high precision like policy forecasting or financial decision-making, 15% could be fatal.

There is another, more fundamental problem: context rot. As the simulation rounds increase, the agents’ conversations get longer, and the model will gradually “forget” early contextual information. MiroFish uses Zep Cloud for long-term memory management, which is currently a decent solution on the market. Zep’s temporal knowledge graph is quite unique among agent memory services. But the README honestly mentions—”Note that token consumption is high, it is recommended to start with simulation attempts of fewer than 40 rounds.” 40 rounds… to be honest, for simulating a complex social system, this ceiling is still too low.

Comparing with Peers

Let’s do a horizontal comparison. Currently, the landscape of LLM-driven social simulations roughly has these tiers:

Academic Benchmark: Stanford’s Generative Agents (now commercialized as Simile). It has papers, benchmarks, validation data, and a $100 million funding round, with investments from Fei-Fei Li and Andrej Karpathy. This currently holds the highest academic credit.

Base Framework: CAMEL-AI’s OASIS. Open-sourced, 2.5k stars, focusing on social platform simulation with million-level agent scalability. MiroFish’s simulation engine is built on this.

Application Layer Product: MiroFish itself. 5.4k stars, 604 forks. Its positioning is to package underlying capabilities into a “prediction tool that ordinary people can use.” Upload materials, describe your needs, and get a report—this interaction design is indeed much friendlier than using OASIS directly.

But there is an issue you need to know: MiroFish’s recommended LLM is qwen-plus from Alibaba’s Bailian platform, and the API call cost for running a complete simulation is not cheap. I haven’t run a full stress test myself, but according to shared experiences, running a simulation with a few hundred agents over a few dozen rounds can easily push token consumption into the millions. If you run it with GPT-4 class models, the cost would be even more exaggerated. This is a common disease of all multi-agent simulation projects—it is essentially a token incinerator.

Also, to be blunt, the current Demo cases—public opinion deduction for Wuhan University and predicting the ending of Dream of the Red Chamber—feel more like “interesting demonstrations” rather than “reliable prediction tools.” The accuracy of public opinion deduction is hard to verify after the fact, and as for Dream of the Red Chamber… who can prove your predicted ending is right? These two cases were chosen cleverly because they evade the most awkward issue of “falsifiability.”

Screenshot of MiroFish project homepage
MiroFish’s homepage style: a fish swimming towards a school of fish in the mirror—”a swarm intelligence mirror reflecting reality.” This metaphor is quite fitting.

Some Things I Sometimes Think About

If we put aside the question of “whether it can accurately predict the future”—a question that may never be perfectly answered—what makes MiroFish genuinely interesting to me is another thing: it lowers the threshold for “thought experiments.”

In the past, if you wanted to run a deduction on “how different groups of people would react if a certain policy were introduced,” you either had to rely on mental visualization or write a paper to build a model. Now, you can throw the relevant materials in and let hundreds of AI characters “play it out” for you. The result may not be perfectly accurate, but it can provide angles you might never have thought of. In this sense, it’s more of a “brainstorming tool” than a “prediction tool.”

I sometimes wonder, what if we used it in product requirement discussions? For example, if you’re building a new feature, instead of finding ten people for a focus group, why not let a hundred AI users “use” it first and see what unexpected behavior patterns emerge? Of course, there’s a paradox here—AI-simulated user behavior is based on training data, and training data is from the past, whereas what you want to predict is the future. Or maybe I’m overthinking it, as human behavior patterns themselves are highly repetitive in many cases.

Another point that concerns me is the project’s AGPL-3.0 license. This means any derivative works based on MiroFish must be open-sourced. For a tool that claims to be a “rehearsal laboratory for decision-makers,” this license choice is… how should I put it, likely to make some enterprise users hesitate. Shanda invested 30 million for commercialization; they will definitely have to find a way to deal with the AGPL restriction down the line.

By the way, I noticed a “cursoragent”—Cursor Agent—in the contributor list. This basically solidifies the Vibe Coding development method: AI writing code for AI to let AI simulate human behavior. The “matryoshka doll” vibes are maxed out.


Having said all this, I’m actually really looking forward to seeing someone use MiroFish to run a prediction with a clear validation window—like the public opinion trend of a specific event next month—and then come back a month later to compare. If anyone is already doing this, I’d really love to see the results.

—— Lyra Celest @ Turbulence τ.


References:

—— Lyra Celest @ Turbulence τ.

Leave a Reply

Your email address will not be published. Required fields are marked *