Two years after its founding by Cambridge neuroscientists who looked to the brain's own distributed architecture for inspiration, London-based Kalosm has drawn $100 million in seed funding — including a $5 million contribution from South Korea's Dunamu & Partners — to pursue a quieter but consequential question: what if the most expensive operation in modern AI has been done wrong all along? By routing fragmented tasks to whichever model and chip combination handles them most efficiently, Kalosm is wagering that intelligence, like the mind itself, is best served not by brute uniformity but by
Dunamu & Partners Backs AI Inference Startup Kalosm in $100M Seed Round
Breaking tasks into smaller units assigned to different models based on efficiency
Why did two neuroscientists decide to start an AI company? That's an unusual path.
They weren't trying to build a better model. They were thinking about how intelligence actually works—how the brain combines specialized circuits rather than just scaling one thing up. They saw the same pattern in AI and realized the industry was doing the opposite.
So they're saying the current approach is fundamentally wrong?
Not wrong, exactly. Single-model inference works fine. But it's not efficient. It's like using a sledgehammer for every task when you have a toolbox available.
The 70 percent cost reduction—is that real, or is it cherry-picked data?
It's from actual financial sector work, not a lab demo. That's the kind of task where speed and reliability matter, so there's real pressure to verify the numbers.
Who else is betting on this idea?
The UK government's sovereign AI fund made Kalosm its first equity investment. That's not casual money. And they're partnering with actual semiconductor companies, not just talking about it.
What happens if this works at scale?
Inference becomes cheaper and faster across the board. That changes what's economically viable to build with AI. A lot of applications that were too expensive suddenly become feasible.
Le Pouls
- Inference costs have quietly become one of AI's most serious bottlenecks, and Kalosm is positioning itself as the answer the industry didn't know it was waiting for.
- A $100M seed round — unusually large, backed by Atomico, the UK sovereign AI fund, and investors from South Korea to Silicon Valley — signals that this is no longer a fringe idea.
- The company's multi-model routing system breaks AI tasks into discrete units and dispatches each to the most efficient model-chip pairing, a direct challenge to the dominant single-model-on-one-GPU paradigm.
- Real-world results from financial sector deployments show 4x faster processing, 70% lower compute costs, and a 10% improvement in task success rates — numbers that move budget conversations.
- With semiconductor partners including Cerebras, Rebellions, and Axelera AI already in the ecosystem, Kalosm is building infrastructure gravity around its approach before competitors can respond.
Two years after its founding by Cambridge neuroscientists who looked to the brain's own distributed architecture for inspiration, London-based Kalosm has drawn $100 million in seed funding — including a $5 million contribution from South Korea's Dunamu & Partners — to pursue a quieter but consequential question: what if the most expensive operation in modern AI has been done wrong all along? By routing fragmented tasks to whichever model and chip combination handles them most efficiently, Kalosm is wagering that intelligence, like the mind itself, is best served not by brute uniformity but by purposeful specialization.
Kalosm, a London startup barely two years old, was built on a neuroscientific intuition: the brain doesn't generate intelligence by scaling a single type of neuron indefinitely — it combines specialized circuits, each suited to different work. Founders Danyal Akarca and Yasha Achterberg, both former Cambridge neuroscientists, applied that logic to AI inference, the computationally intensive process of running trained models to produce outputs, and asked whether the industry's default strategy of routing everything through one massive model on one type of chip was truly optimal.
Their answer, backed now by $100 million in seed funding, is that it isn't. The round was led by Atomico and included Plural, DCVC, and a symbolically significant participant: the UK government's sovereign AI fund, which chose Kalosm as the target of its first equity investment. Dunamu & Partners, a South Korean firm, contributed $5 million — roughly 7 billion won — reflecting the cross-geographic appetite for what Kalosm is building.
The system works by decomposing AI tasks into smaller units and routing each to whichever model and semiconductor combination handles it most efficiently, optimizing across cost, power, and speed. Kalosm is not building this in isolation; partnerships with Cerebras, Rebellions, Axelera AI, Tengr.ai, and Supermicro are constructing an ecosystem around the approach.
The numbers from real-world financial sector deployments are difficult to dismiss: processing speeds four times faster, computing costs reduced by 70 percent, and task success rates up 10 percent compared to conventional GPU-based single-model inference. For organizations where inference costs are already constraining AI deployment at scale, those figures reframe the conversation entirely. Whether Kalosm's distributed inference architecture becomes a niche tool or rewrites how the industry thinks about running AI will be one of the more consequential questions of the next few years.
Kalosm, a London-based startup founded just two years ago by two former Cambridge neuroscientists, has closed a $100 million seed funding round—a striking vote of confidence in an approach to artificial intelligence that challenges how the industry currently runs its most expensive operations. Danyal Akarca and Yasha Achterberg built the company on an observation drawn from neuroscience: the brain doesn't generate intelligence by endlessly multiplying neurons of a single type. Instead, it combines circuits with different functions, each specialized for different tasks. They applied that insight to AI inference—the computationally intensive work of running trained models to produce outputs—and asked whether the industry's dominant strategy of running a single massive model on one type of chip was actually the most efficient path forward.
The funding round, led by European venture capital firm Atomico, included backing from Plural and DCVC, along with a notable institutional player: the UK government's sovereign AI fund, which selected Kalosm as the target of its first equity investment. Dunamu & Partners, a South Korean investment firm, contributed $5 million to the round, translating to roughly 7 billion won. The capital influx signals that investors across geographies see real potential in Kalosm's core thesis.
That thesis is straightforward in concept but demanding in execution. Rather than sending every task through a single large model running on a single type of semiconductor, Kalosm's system breaks work into smaller, discrete units and routes each to whichever model and chip combination best handles it—optimizing for cost, power consumption, and speed depending on what the task actually requires. The company isn't working in isolation. It's partnering with semiconductor and infrastructure firms including Cerebras, Rebellions, Axelera AI, Tengr.ai, and Supermicro, building an ecosystem around the multi-model approach.
The real test of any infrastructure technology is whether it delivers measurable gains in the real world. Kalosm has published results from financial sector AI agent tasks—work that typically demands both speed and reliability. Compared to the conventional single-model approach using existing GPU infrastructure, Kalosm's method achieved processing speeds four times faster. Computing costs dropped by 70 percent. Task success rates improved by 10 percent. Those aren't marginal improvements. They're the kind of numbers that make CFOs and infrastructure teams sit up and pay attention, especially in an industry where inference costs are becoming a serious constraint on scaling AI applications.
The timing matters. As AI models have grown larger and more capable, the cost of running them has become a bottleneck for many organizations. A 70 percent reduction in inference costs could reshape the economics of AI deployment across industries. The speed gains are equally significant for applications where latency matters—financial trading, real-time decision-making, interactive systems. Kalosm's approach suggests that the next wave of AI efficiency gains may come not from building bigger single models, but from getting smarter about how work is distributed across the tools already available. The company's success in the coming years will likely determine whether this becomes a niche optimization or a fundamental shift in how the industry thinks about inference.
Citations marquantes
The brain creates intelligence by combining circuits with different functions rather than endlessly increasing one type of neuron— Kalosm founders' founding principle