News

Apigee adds model and MCP tool performance dashboards

In brief

Apigee’s new AI and MCP dashboards separate model and tool traffic. Learn which filters and proxy policies make their measurements useful.

3 min read

Sources
Editorial illustration of three swimmers in separate pool lanes, each at a different position.
Original AI-generated editorial illustration of separate measurement populations; not a product interface.
On this page2 sections

When an incident assistant takes too long to answer, the model is only one place the time can go. It may also be waiting for a tool to retrieve logs or deployment history. Google Cloud added two API insights dashboards on September 30: AI performance for model traffic and Tool performance for Model Context Protocol (MCP) tools. For teams already routing this work through Apigee, they offer a more direct way to examine those two parts of an agent’s work. The change is listed in Google’s release notes.

The useful question is which part of the request became slower. A model change, a busier log-search tool and a shift toward longer investigations can all make the overall experience worse, but they suggest different next checks. These dashboards supply measurements for that investigation; interpreting them still depends on selecting the traffic that matches the complaint.

Model calls and tool calls get separate views

The AI view reports token consumption, model latency and errors. The Tool view reports MCP traffic, throughput, payload sizes and latency, with filters including tool name and deployment. In the console, both sit under API hub → API insights. Google’s dashboard reference contains an easily missed detail: latency percentiles pool models or tools until a specific model or tool filter is applied.

Consider a hypothetical incident assistant using two tools: get_deployment retrieves one deployment record, while search_logs scans a larger body of events. If users start asking more log-heavy questions, the combined latency distribution can change even if neither tool has slowed down. Conversely, a busy, fast deployment lookup can make a combined view less representative of the slower log searches that responders are waiting for.

Start with the affected tool and the same time window as the reported delay, then carry that selection into the backing service’s logs or traces. This preserves the question as the investigation moves between screens. Comparing an incident’s log searches with an all-day mixture of unrelated calls would change the population halfway through the diagnosis.

An incident assistant calls a model and two different tools. Select search_logs and the affected time window before following that tool into backend evidence.
Illustrative investigation path, not an Apigee screen. Keep the tool and time window consistent when moving from aggregate measurements to backend evidence.

The charts need the right proxy policies

Google requires VerifyAPIKey, PromptTokenLimit and LLMTokenQuota on the LLM-facing proxy to populate all AI charts and filters. Without API-key verification, total tokens can appear while the application breakdown stays empty. A partly populated dashboard can therefore indicate incomplete instrumentation rather than an absence of application traffic.

Configuration deserves a separate check from visual appearance. Google’s token-policy tutorial distinguishes counting tokens from enforcing quotas: CountOnly records usage, while EnforceOnly can reject requests. It also requires explicit JSON paths for the relevant prompt and response fields. A monitoring trial should not accidentally introduce a new rejection policy simply because a sample configuration includes one. The tutorial’s token-policy guidance applies to Apigee, not Apigee hybrid; confirm the supported configuration for your deployment before following it.

For an existing Apigee user, a useful first check is one known model request and one known tool call, followed through the relevant filters. Verify that the expected application, model and tool appear, then compare the displayed interval with the original request records. The broader dashboard investigation guide explains why preserving that context matters across monitoring tools.

If the calls bypass these gateways, this view cannot account for their delay. Use the application’s tracing for that path rather than concluding that an empty gateway chart clears the dependency. Where the traffic is present, the new dashboards can shorten the route from “the assistant is slow” to a specific model or tool worth investigating, without requiring that first question to become a custom reporting project.

Sources & context

Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.

Report an error or outdated detail

Related reading

Explore a related question