AI Monitoring

On the [AI Asset > AI Monitoring] page you review how your monitored AI Assets are actually behaving — what looked unexpected, what it cost, and how many tokens it was consumed by. The assets and the detection rules behind these figures are registered under AI Management.

Tabs

AI Monitoring holds three tabs over the same set of assets, each answering a different question.

TabQuestion it answers
AnomaliesDid anything behave unexpectedly?
AI CostWhat is the usage costing?
AI UsageHow many tokens are being consumed, and by which model?

Each tab keeps its own filters in the URL, so switching tabs does not carry one tab’s filters into another.

Anomalies

The Anomalies tab collects every usage anomaly detected across your monitored AI Assets, so you can triage them in one place.

ℹ️
This board covers AI assets only. Anomalies on ordinary cloud resources are raised by a different feature and are read under Resource Monitoring — the two boards never mix, so a count here is a count of AI anomalies.

Situation Card

The card above the table is a status line, not a metric.

StateMeaning
All clearNothing is currently firing
n firingAnomalies are in progress. Click the card to filter the table down to just those

Next to it, Anomalies · last 14 days charts daily counts so a quiet week and a noisy one are distinguishable at a glance, with today’s count called out on the right.

Viewing

Search matches the model. The dropdown filters by status, and the ⋮ button opens column and page-size settings.

ColumnDescription
StatusWhether the anomaly is still firing or has resolved
ProviderCloud provider of the asset
ModelThe model the anomaly was detected on, with its AI service and token direction (input / output)
AccountCloud account the asset belongs to, named the same way the cloud account list names it
AppApplication the asset is assigned to, by name, or Unassigned
DetectionThe rule that flagged it, written as metric · sensitivity — for example Token usage · Medium
DeviationHow far the measured value strayed from its expected range
TrendA sparkline of usage around the event
First seen · ResolvedWhen the anomaly was first seen and when it cleared, with how long it lasted

A summary line under the table reports the total split into firing and resolved.

ℹ️
The list is ordered by when an anomaly was first seen, not by when it was raised. A condition can hold for a while before the rule trips, so the two differ — ordering by first sight keeps an event in the place where it actually began. The detail panel still shows both moments.

The leading columns stay in place while you scroll the table sideways, so the model and status do not scroll away from the numbers.

Sensitivity

Rules are tuned by sensitivity rather than by a hand-set number. The value trades misses against false alarms:

SensitivityBehaviour
HighFlags a value that strays only slightly outside its usual range. Fewer misses, but more false alarms
MediumThe default. Balances misses and false alarms for most metrics
LowFlags only a value far outside its usual range. Fewer false alarms, but small anomalies may slip through
ℹ️
If a rule shows Unknown, the server returned a sensitivity this version does not recognise — check the rule in AI Management.

Anomaly Detail

Clicking a row opens the event. The title is the model, with the status beside it and the metric, account, App, and AI service underneath.

Detection evidence

The chart plots the usage time series around the event, with the anomaly window shaded and three moments marked:

MarkerMeaning
OnsetWhen the condition started
FiredWhen it had held long enough to be raised as an anomaly — the delay between Onset and Fired is the rule’s Sustain
ResolvedWhen usage returned to normal

Where the rule has a threshold to draw, a single Threshold line is laid over the series so the gap between usage and the line is readable at a glance.

ℹ️

Stretches where no baseline had been established are drawn as a distinct No baseline region rather than as a flat line at zero — the evaluation did not run there, which is different from “usage was zero”. Hovering such a point says No evaluation ran at this time, and a point carried over from the previous evaluation is marked Backfilled.

The region runs right up to where the normal-range band starts, so there is no unexplained gap between “nothing was being judged” and “judging began”.

⚠️
When the threshold and the usage differ by orders of magnitude the y-axis switches to a log scale, and the chart says so. Height differences on a log axis are not proportional — read the numbers, not the bar heights.

Summary

The right-hand panel restates the event as a sentence — what rose, how far, and whether it is still going — then backs it with the numbers. It is written in the terms the screen uses elsewhere (actual usage, expected usage, normal range) rather than in statistical notation:

ItemMeaning
DeviationHow far the value strayed from what was expected
Actual usage / Expected usageThe measured value the verdict was made on, against what was expected of it
Normal range upperThe top of the normal range — the value being over this is what tripped the rule
Window / SustainThe rule’s observation window and how long the condition had to hold
First seen / Fired / Resolved / DurationThe event’s timeline

View in AI Management opens the asset behind the anomaly and Edit Rule jumps to the rule that raised it. Under Related views, This model’s usage and This model’s cost carry the same model into the AI Usage and AI Cost tabs.

ℹ️
A rule can apply to several anomalies at once. When it does, the panel says so — editing that rule changes detection for the others too, not just this event.
ℹ️
With no anomalies detected, the situation card reads All clear and the table is empty. That is the normal state — anomalies only appear once a monitoring rule on a registered asset trips. If the table stays empty when you expect events, check on AI Management that the assets are switched on and have active rules.

AI Cost

The AI Cost tab breaks down what your monitored AI usage actually costs, split into what has already been billed and what is still an estimate.

ℹ️
Cloud providers publish billing on a delay, so the most recent days have no billed figure yet. Rather than showing a gap, this page estimates the unbilled tail and marks it as such everywhere it appears — the boundary between the two is drawn on every chart.

Filters

FilterChoicesWhat it changes
ServiceThe AI service to analyzeScopes the whole page to one service
Cost BasisEffective Cost · Billed Cost · List CostWhich price the numbers are based on
GranularityDaily · MonthlyThe bucket size of the main chart
PeriodStart and end dateThe query range
Group byModel · AssetWhether the chart and detail table split by model or by asset
FilterAsset · Model · RegionOptional narrowing on top of the above

A line above the card reports how many assets and models are in scope, so you can tell at a glance whether a filter cut more than you intended.

Summary Cards

CardMeaning
Total costCost over the selected period. When part of it is estimated, a caption says how much
Month-end forecastProjected cost for the full month, with the current daily average underneath
Cost compositionA bar splitting the total into Billed and Estimated, with both amounts

Cost composition is the quickest way to judge how much of what you are looking at is settled. A total that is mostly estimated will still move as billing catches up.

Model — cost & tokens

Bars per period, coloured by model (or by asset, if you grouped that way). The subtitle states the cost basis and the date billing runs through, and a shaded region marks everything past the Billed / Estimated boundary.

The Cost / Tokens toggle switches the same bars between money and token counts, so a model that is cheap but chatty is easy to spot.

ℹ️
The chart also marks Monitoring starts — the date CloudOps began measuring tokens for this asset. Before it there are billed charges but no token measurements, so the caption underneath says so directly rather than letting the empty stretch read as zero usage.
⚠️
The Estimated rate chip is not a published provider price. It is simply the billed amount divided by the tokens billed in the same period. Within a single model the real price differs by how the tokens were used — input, output, caching are each priced differently — so treat this as a blended average, not a rate card.

Detail by Model

ColumnDescription
ModelModel name, colour-matched to the chart
Billed costCost the provider has already billed
Est. costEstimated cost for the not-yet-billed part
Est. rate /1MThe blended rate described above, per million tokens
Billed tokensTokens the provider billed for
Metered tokensTokens CloudOps measured itself
DiffThe gap between the two token counts

The Usage link on each row carries that model into the AI Usage tab.

ℹ️
Billed tokens and Metered tokens are counted by two different parties — the provider’s billing and CloudOps’ own metering — so Diff is a sanity check rather than an error. A dash means there is nothing to compare in the selected range.
⚠️
The cost columns and the token columns do not cover the same span. Costs are totalled over the whole query range, while tokens are compared only over the range where both billed and metered figures exist. A caption above the table states both spans — read it before comparing a cost against a token count in the same row.

When there is no measurement

Billed amounts and metered tokens do not always start on the same day, and the screen says so instead of showing a zero:

SituationWhat the screen says
Metering began partway through the rangeToken measurement starts <date>. Before that, only billed amounts exist — no measurement.
The asset is registered but nothing has been metered yetRegistered <date>, but no token measurement has arrived yet. Only billed amounts are shown.
A single row has no metered figureThe token cell reads Not metered
ℹ️
Not metered is not the same as zero usage. It means CloudOps has no measurement for that span — the billed amount beside it is still real.
⚠️
An Unmapped cost warning means a new SKU appeared that is not yet in the rate lookup, so its cost could not be attributed to a model. The amount is real but unattributed, and needs review.

Month-to-date Trend

Cumulative spend from the first of the month, always for the current month regardless of the Period filter above.

LineMeaning
This month billedSolid — the settled figure, up to the billed boundary
With estimateDashed — the same curve continued with estimated cost
Prev month, same periodLast month at the same day-of-month, for comparison

Two vertical markers frame the reading: Billed boundary where billing data stops, and Today. The region past today is shaded, since nothing there has happened yet. Under the chart, the cumulative figure and the current daily average are the two numbers behind the Month-end forecast card.

AI Usage

The AI Usage tab breaks down one asset’s metered token usage by model and by token type, at minute, hour, or day resolution. Where AI Cost answers what is this costing, this tab answers what is actually being consumed.

ℹ️
Everything on this tab is metered — measured by CloudOps directly, not taken from provider billing. That is why the numbers here can be read minutes after the usage happens, while billed cost lags by days.

Filters

FilterChoicesWhat it changes
Service / AssetThe provider service and one registered AI assetEverything on the tab is scoped to this asset
ModelAll models, or oneNarrows the chart and table to a single model
RangeMinute · Hour · Day, plus a length slider (7d / 14d / 30d / 90d)Both the bucket size and how far back to look
ℹ️
This tab looks at one asset at a time, unlike the Anomalies and AI Cost tabs which span everything you monitor. Resolution and length are picked together: minute resolution over a long range would be unreadable, so the slider offers the lengths that make sense for the resolution you chose.

Summary Cards

CardMeaning
Total tokensInput plus output over the selected range
Input tokensTokens sent to the model
Output tokensTokens the model generated

Keeping input and output apart matters because they are priced differently — a workload heavy on output costs more than the same token count spent on input.

Provisioned Throughput

Where the asset’s account holds Vertex AI provisioned throughput, a snapshot of that reservation sits above the token totals.

CardShows
UtilizationHow much of the reserved capacity is actually being used, with a status badge
ProvisionedThe capacity you bought, as a token limit in tok/s
Consumed nowWhat is being drawn right now — dedicated capacity only, with on-demand overflow excluded

The status badge reads Underused, Healthy, Scale up, or Awaiting data, and Over reservation when consumption has passed the limit.

⚠️
Provisioned throughput is a commitment: unused capacity does not roll over and it cannot be cancelled mid-term. Review the commitment when utilization stays low, and scale up when it runs high.
ℹ️
On an account with no provisioned throughput the whole section is absent — not shown as 0%. An empty card would read as “we bought capacity and are not using it”, which is the opposite of the truth.

Token Usage

Stacked bars per bucket, at the resolution set in Range. The toggle on the right switches how the stack is split:

ViewSplits the bars by
Input/OutputToken type — the shape of the workload
By modelModel — which model is doing the work

Use Input/Output to see how the models are being used, and By model to see which model to look at next.

Detail by Model

ColumnDescription
ModelModel name, colour-matched to the chart
Input tokensTokens sent to that model over the range
Output tokensTokens it generated

The subtitle repeats the point worth remembering: these are metered figures, neither billed nor estimated. To see the same models expressed as money, switch to the AI Cost tab.

ℹ️
With no usage in the selected range, the chart is empty and the cards read zero. Widen the range with the slider, or check on AI Management that the asset is switched on — a paused asset stops being metered.
v1.7.0