Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Latest AI and industry news. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

News AI News

A controlled test says the expensive "skill knowledge graph" layer for AI agents is worse than plain search

Researchers built a typed knowledge graph (1,421 AI-generated connections, ~$2.70 to generate) on top of a 690-skill library to help an agent pick the.

Enterprise DNA |
A controlled test says the expensive "skill knowledge graph" layer for AI agents is worse than plain search

AI Pulse · Under the Radar

The play

Skip the fancy knowledge graph layer for now, plain keyword search on your skill library will outperform it and cost almost nothing.

A team ran a controlled experiment on something a lot of AI vendors are quietly betting on: building a fancy knowledge graph to help agents find the right skill. The idea sounds sensible. You tag 690 skills with typed relationships, spend a few dollars generating 1,421 connections, and the agent should navigate better than dumb keyword search.

It didn’t work. At the same token budget, the knowledge graph scored 0.632 while plain search hit 0.744. Worse, the same system scored 95% on one set of test questions and 74% on another set covering identical tasks, just written by different people. That variance alone should make anyone nervous about production reliability.

This isn’t a fluke. The same week saw a cluster of concurrent arXiv preprints all wrestling with agent skill retrieval, which means multiple labs hit the same wall at the same time. The pattern matters more than any single paper.

What it means for your setup

If you’re building or buying an agent system, ask hard questions about how it picks which tool to use. A lot of pitches will talk about semantic layers, ontologies, or graph structures that sound impressive. This test suggests the simpler approach works better and costs less to run.

For anyone running something like a skill registry, whether it’s EDNA’s .claude/skills/ folder or your own internal library, the takeaway is blunt: don’t over-engineer retrieval until you’ve proven plain search fails. We see this in the Omni Command Centre design all the time. Teams want the sophisticated answer first, but the boring one usually ships faster and breaks less.

The $2.70 generation cost is trivial. The reliability gap and the question-sensitivity aren’t. If your agent can’t consistently find the right skill because someone phrased a request differently, you have a production problem no graph will fix.

Free daily email

Get this every morning.

This brief is one item from today's AI Pulse, the short daily read we run for ourselves on what is actually happening in AI. Subscribe free and it lands in your inbox each morning.

Free daily email

Subscribe to the daily AI Pulse

One short read every morning on what is actually happening in AI. Free.

One email a day. Unsubscribe any time.