Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Latest AI and industry news. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

News AI News

NVIDIA publishes the first controlled A/B numbers on "skills" moving outcomes

A run across 300+ of its verified agent skills, isolating the skill as the only variable, reports +41 points correctness and +39 points effectiveness..

Enterprise DNA |
NVIDIA publishes the first controlled A/B numbers on "skills" moving outcomes

AI Pulse · AI Trends Pulse

The play

Add validated skills to agent workflows experimentally, measuring correctness and effectiveness against your current baseline before broad rollout.

NVIDIA ran a controlled test on its own agent skills, the add on capabilities that plug into an AI agent to make it better at a specific task. They took over 300 verified skills and isolated the skill itself as the only thing changing between runs. The result was a jump of 41 points in correctness and 39 points in effectiveness. That’s a big claim, and it’s worth saying plainly that this is vendor published data, so nobody outside NVIDIA has checked it yet. Read the original post and treat the numbers as a starting point, not a verdict.

Here’s why this matters even with that caveat. Most business owners assume an AI agent is either good or bad, like it’s one fixed thing. What this test suggests is that a big chunk of performance comes from the skills bolted onto the agent, not just the underlying model. If that holds up under independent testing, it changes how you should think about buying or building AI tools. The question shifts from “which model is best” to “which skills does this thing actually have, and can I swap in better ones.”

For a real business, that’s a practical shift. It means the AI setup you have today isn’t locked in. If a specific skill, say pulling accurate numbers from a spreadsheet or writing a compliant customer email, is weak, you may be able to fix that one piece without replacing the whole system. That’s the exact idea behind giving your team a proper AI command centre, where you can see which skills are actually driving results and swap out the weak ones, and it’s the kind of thing we build into an AI command centre. Worth watching for replication before you take the numbers as gospel, but the direction is worth paying attention to.

Working With Claude field guide cover

Free Resource

Put what you just read to work

The free 32-page Working With Claude guide: the full ecosystem, Claude Code, and how to roll it out across a business.

No spam. Unsubscribe any time.

Free daily email

Get this every morning.

This brief is one item from today's AI Pulse, the short daily read we run for ourselves on what is actually happening in AI. Subscribe free and it lands in your inbox each morning.

Free daily email

Subscribe to the daily AI Pulse

One short read every morning on what is actually happening in AI. Free.

One email a day. Unsubscribe any time.