Deepgram has added two observability options for customers running its speech-to-text and text-to-speech models on Amazon SageMaker AI. The update exposes information that was previously held inside the vendor container, including the usage values behind marketplace billing and the behavior of the inference engine on each GPU.
Enhanced Metrics sends usage, billing and feature-level measurements directly to the customer’s Amazon CloudWatch account. Deepgram says it requires no separate agent, sidecar container or additional identity permissions. The values match the consumed units used for AWS Marketplace metering, allowing teams to compare their bill with traffic by model and transport.
The deployment also supports Prometheus and OpenTelemetry collection for engine, accelerator and host metrics. This matters for capacity planning because an endpoint being online says little about why a GPU is saturated or what traffic is driving cost. Audio and transcripts remain inside the customer’s AWS account, although organizations still need to assess their own retention, access and compliance controls.