Intelligence, the startup behind the AI design evaluation platform DesignArena, announced Monday that it has raised $7.9 million in seed funding. The round was led by Index Ventures, with participation from Conviction, A*, Valkyrie, and others. The company says the platform, which uses crowdsourced human rankings to help AI models improve their output, is now used by 5.3 million people worldwide and generates $60 million in annual recurring revenue.
The company traces its origins to a college project in 2025, when co-founder Grace Li and a few friends were building an AI game engine. The models could produce functional games, but none of them were fun — leading the group to a fundamental question: how do you tell if a game will be fun? They concluded there was no substitute for human judgment, and began exploring ways to gather honest feedback at scale.
Also read: ChatGPT dominates Congress's AI spending, capturing roughly 90% of tool purchases
That experiment became DesignArena. For individual users, the platform works like a sophisticated model router. Users type a prompt into a ChatGPT-style window, choose a format — websites, images, and a dozen other visual categories — and then rank the resulting outputs through a series of A/B comparisons. The process yields a ranked list from best to worst, which is useful on its own. But the real value, Li says, lies in the enterprise side: AI labs pay to access this constant stream of human preference data to refine their media-generating models.
“It was the missing bottleneck for a lot of these models to make improvements in the design space,” Li said. “About a week later, we closed our first major deal with a frontier lab, and the rest is kind of history.”
Also read: Sam Altman calls for 'pacing' AI development after OpenAI agent breached Hugging Face
Why human feedback is suddenly valuable
DesignArena’s growth comes as AI companies grapple with the limits of automated benchmarks. These tests can run at massive scale, but they are often vulnerable to gaming. The recent breach at Hugging Face, which exposed vulnerabilities in automated evaluation systems, highlighted this risk. Crowdsourced human evaluation offers a complement — it is slower and more expensive, but it captures qualitative judgments that automated metrics miss.
Li notes that user preferences vary by region and over time. For instance, web dashboards in Asia tend to favor a more maximalist design style. By requiring users to log in to receive their outputs, Intelligence can track these shifts across continents, giving frontier labs a richer picture of what users actually want.
The market for human-led evaluation is not without its cautionary tales. Yupp, a similar startup that raised $33 million from investors including a16z crypto’s Chris Dixon, shut down earlier this year after less than a year in operation. It had attracted some frontier model customers and claimed over 1.3 million users, but could not build a sustainable business. That failure underscores the challenge of converting user engagement into a durable revenue model.
What the funding means for the AI evaluation space
Despite Yupp’s collapse, other players are thriving. LM Arena, which applies a similar approach to text-based responses, raised $150 million in a Series A in January, just four months after launching its paid product. The contrast suggests that while the market is real, execution and timing matter enormously.
For Intelligence, the new capital will be used to expand its enterprise offerings and deepen its data analytics capabilities. The company is positioning itself as a critical piece of infrastructure for AI development — a source of human judgment that can help models become not just more capable, but more aligned with what people actually want.
The broader implication is that as AI models become more sophisticated, the bottleneck may shift from raw compute to data quality. DesignArena and its peers are betting that human taste, in all its subjectivity, is a valuable commodity. Whether that bet pays off will depend on whether they can maintain user engagement and prove their data’s worth to the labs that need it most.
For now, the company’s rapid growth — from a college project to a $60 million ARR business in about a year — suggests there is genuine demand. The next test will be whether Intelligence can scale without losing the community that makes its data valuable.
Disclaimer: This article is for informational purposes only and does not constitute financial advice. The markets and technologies discussed are volatile and subject to rapid change; readers should conduct their own research before making any investment decisions.