As AI workloads strain computing budgets, companies are looking for ways to accelerate token generation without buying entirely new hardware architectures. Kog’s approach targets mainstream chips like the AMD MI300X and Nvidia H200, promising up to 30x faster decoding speeds by reverse-engineering hardware at a low level to tap into unused memory bandwidth.
Proving the Concept on Standard GPUs
Kog first grabbed attention in May with a tech preview demonstrating that extremely fast single-request decoding is possible on standard enterprise hardware. The demo achieved an impressive 3,000 per-request tokens per second (TPS) using an open-sourced 2-billion parameter model called Laneformer 2B.
The performance struck a chord. CEO Gaël Delalleau noted that the company received 200 tangible business leads following the preview. Early feedback points to software engineering as the primary use case, particularly for developers frustrated by slow AI coding assistants. Anthropic already charges a premium for faster Claude responses, highlighting the market’s willingness to pay for speed. Kog is targeting customers who rely on these AI workflows for professional tasks and cannot afford hours of waiting for results.
Generating games and applications via prompt is another design partner use case. For these creators, a faster outcome thanks to the Kog Inference Engine (KIE) directly translates to increased revenue, as quicker iterations allow for more product development cycles.
The Challenge of Scaling LLM Inference
Despite the early success with small models, Kog faces a massive hurdle scaling its technology to the large language models (LLMs) that enterprises actually want to use. Delalleau admitted that prospective customers are not prepared to fine-tune smaller models to meet their needs. In response, Kog has shifted its focus entirely to accelerating larger models to meet actual market demand.
This requires a meticulous, hardware-level approach. Delalleau brings a unique background to the problem, combining studies in solid-state physics at France’s École Polytechnique with experience in offensive cybersecurity. A four-time finalist at DEFCON’s Capture the Flag tournament, he applies a hacker mindset to GPU engineering. He believes that the idea standard GPUs are poorly suited for decoding is a misconception, arguing that newer chips possess memory bandwidth that simply begs to be unlocked.
The downside of this reverse-engineering approach is that it is highly time-consuming. Delalleau noted that for every new GPU, the team dedicates weeks or even months to digging into the details. This severely limits the number of chips Kog can support in the near term, making strategic hardware choices vital.
A Crowded AI Acceleration Market
Kog is not alone in the race to accelerate AI. The market recently welcomed Cerebras and its purpose-built chips to public markets, while other French startups like ZML are developing hardware-agnostic software to bypass Nvidia’s CUDA. However, Kog likens its deep-level focus to Stanford University’s Hazy Research lab, aiming to extract maximum performance from existing silicon rather than replacing it.
This software-first strategy offers a distinct advantage for enterprises heavily invested in current hardware. While specialized chips promise breakthrough speeds, the capital expenditure required to overhaul datacenter infrastructure remains a significant barrier. Optimizing existing datacenter GPUs provides a practical bridge for companies looking to improve AI performance immediately without abandoning their hardware investments.
What Happens Next
Kog must prove its 30x claim works on major LLMs, not just 2-billion parameter models. The startup plans to implement its first major model at 10x speed by September. Achieving this milestone will be critical for demonstrating customer traction and securing a Series A funding round.
Looking further ahead, Kog hopes to automate its manual optimization process using agent-based pipelines, allowing the team to support a wider variety of chips and models without the intense time commitment. As Europe pushes for greater technological sovereignty in AI hardware and software, Kog’s France-based development—backed by Scaleway, Bpifrance, and French Tech 2030—could position the startup as a key player in the region’s AI acceleration ecosystem.
— David Kim, technology desk, AXO News