Infinity Raises $15M for AI Agent That Writes Chip Code, Cracking Nvidia’s CUDA

AI startup Infinity raised $15 million at a $100 million valuation Monday for an AI agent that writes chip kernels, targeting Nvidia's CUDA software moat.

AI-generated Axo News staff avatar for David Kim
5 Min Read

The San Francisco company closed its seed round with backing from Touring Capital, Principal VC, executives at major chip firms, and researchers from OpenAI and Anthropic. Its product, an AI agent called Ignition, generates, tests, and rewrites the low-level code that makes AI models run on specific chips. That work has historically kept developers locked into Nvidia’s hardware and software stack.

Why Nvidia’s CUDA Lock-In Matters

Nvidia holds an estimated 80% of the data-center AI accelerator market, but raw silicon speed is only part of the story. The company’s deeper advantage is CUDA, a software layer built up over nearly two decades that lets PyTorch and TensorFlow run on Nvidia hardware by default. Write a model in Python, and it just works on Nvidia.

Rival chips from AMD, Qualcomm, and AWS routinely match Nvidia on compute power. What they lack is the software ecosystem. Porting a model to new silicon means writing custom kernels, the low-level code that drives a chip. Few engineering teams can afford the months of work that requires. That gap is the CUDA moat.

How Ignition Automates Kernel Generation

Infinity’s Ignition agent attacks that gap directly. It generates kernels, tests them, rewrites them, and keeps tuning as performance data flows back. Human engineers set the direction. The agent handles the grind. The approach mirrors the founder’s broader obsession with what he calls “automated invention.”

The reported results are striking, though they come from Infinity itself. Working with chip maker d-Matrix, Ignition hit 92% of a new chip’s peak performance 10 hours after first touching the hardware. Within 10 days, three frontier models were running on it end to end. On a separate test, Infinity says it lifted a model’s output from about 1,400 to more than 20,000 tokens a second in a single day. That would beat vLLM, a widely used open-source inference framework. No one has independently verified the claims.

The Founder and the Business Model

Jeremy Nixon, a former Google Brain researcher, leads the company. He created AGI House, a San Francisco hacker network that says it has spawned hundreds of startups. Nixon previously built Omega, an algorithm that invented other machine-learning algorithms and scored them in a loop. Ignition applies the same instinct to hardware kernel generation.

Nixon announced the raise on X, hailing the arrival of AI systems that “enable, optimize and invent” the next generation of AI systems. He has framed the stakes in expansive terms. “The next era of AI will be defined not just by who makes the best chip, but by who can make any chip run state-of-the-art models at blazing speeds,” he said.

Infinity says it already earns millions in annual recurring revenue and employs 26 people. Rather than charging a licence fee, it takes a cut of the speed and cost gains it delivers to customers. That aligns its income with measurable performance improvements — a model that could scale quickly if the benchmarks hold up.

A Crowded Race to Break Nvidia’s AI Chip Software Lock-In

Infinity is not alone in targeting Nvidia’s AI chip software advantage. A wave of startups is working to make AI models portable across silicon. They bet that AI inference — projected to become two-thirds of all AI compute spending this year — will reward whoever can run it cheapest and fastest. Cheaper AI inference is worth a fortune as model deployment scales.

The caveats are real. Infinity is a seed-stage firm with one public chip partner, and it self-reports its headline benchmarks. Nvidia, meanwhile, is not standing still. The company continues to invest in CUDA and its growing software stack, and rivals have struggled for years to build a credible alternative.

What Happens Next

Watch for two signals over the next year. First is whether independent benchmarks confirm Ignition’s reported gains, particularly the token-per-second jump against vLLM. Second is whether Infinity signs additional chip partners beyond d-Matrix. That would test whether the Ignition agent generalizes across architectures or works only on a narrow set of hardware.

Nvidia’s software moat has survived many challengers. The question is whether AI-generated kernels finally change the economics of porting models — or whether this seed round becomes another footnote in the long history of attempts to break Nvidia’s grip on AI software.

— David Kim, technology desk, AXO News

Share This Article