French startup Kog is arguing that GPUs still have more room to support agentic AI workloads, TechCrunch reports.

The company’s premise challenges a growing complaint in AI infrastructure: that GPUs, while excellent for dense model computation, are less efficient for agentic workflows that involve many small calls, tool steps, and irregular execution. Kog is working deeper in the stack to squeeze more inference out of existing GPU hardware.

If the approach works, the consequence would be practical rather than flashy. Companies could run more agent activity on the hardware they already buy, reducing pressure to add capacity or move specialized workloads elsewhere. That matters as inference, not only training, becomes a major cost center for AI products.

The report does not mean GPUs are the perfect answer for every agent system. Memory, networking, scheduling, and software overhead still shape real performance. But it shows that the infrastructure debate is not simply GPU versus non-GPU hardware. There is still room for optimization in how inference work is packed and served.