Goodfire has launched monitoring tools that inspect internal activations—the numerical signals inside a model—while an AI agent works. The monitors are available to Baseten customers and initially support the open Kimi K3 model.
Small classifiers called probes run at each step and look for patterns linked to risks such as offensive hacking, chemical or biological misuse and reward hacking. A separate language model reviews only the activity that a probe flags. Customers can log the event, route it to a person or refuse the request.
Goodfire says this is cheaper than asking another large model to reread an agent’s complete trace. In company tests covering about 1,500 Kimi K3 sessions, internal monitoring cost roughly $51, compared with $233 for a cheaper external reviewer and about $10,000 for a top-tier one. The probes detected 94 percent of malicious hacking sessions while escalating 8.7 percent of harmless sessions. Four probes reportedly added less than 2 percent to initial response latency.
Those are vendor-run results on one model and threat set, not independent proof across production workloads. Internal probes may complement sandboxing, permission limits and human review, but they cannot replace controls that prevent an agent from reaching sensitive systems in the first place.