Condition-Based Maintenance Pilot Metrics for Transformer APM
A practical scorecard for measuring a transformer condition-based maintenance pilot across evidence completeness, review speed, decision quality, execution, and governance.

On this page
A condition-based maintenance pilot can fail while the software is functioning. It may have no agreed cohort, no baseline, inconsistent timestamps, unclear review ownership, or a scorecard that jumps straight to “savings” before the team knows whether the evidence is complete. Transformer APM pilots need a narrower first question: can the organization make a more timely, reviewable, and repeatable decision from the evidence it already has?
This is the measurement logic behind the transformer CBM pilot value calculator. The tool is a planning aid. A production pilot should use the utility’s own records, definitions, approval process, and cost model.
Build the baseline before the dashboard
Define the pilot cohort by asset ID, voltage class, component scope, criticality, and evidence availability. Record the observation window, normal sampling cadence, online-monitor coverage, planned outages, existing maintenance strategy, and who reviews alarms today. If the pilot mixes large power transformers with distribution units or combines laboratory-only and online-monitored assets without labeling them, the results will be difficult to interpret.
Create a data dictionary for terms such as “alarm,” “review started,” “review complete,” “approved action,” “urgent work,” “repeat sample,” and “missing evidence.” Capture the source timestamp, not only the time an analyst opened a record. A rate or latency calculated from inconsistent clocks can look precise while being wrong.
Five metric families
1. Evidence readiness
Measure the percentage of cases with a valid asset ID, source record, timestamp, unit, method, monitor status, data-quality flags, and links to supporting history. Track missing DGA sample dates, unverified monitor readings, absent load context, stale inspection notes, and duplicate records. CIGRE TB 761 describes assessment indices and consistent scoring; the pilot should make its inputs inspectable before it ranks assets.
2. Workflow speed
Measure time from source arrival to triage, triage to qualified review, review to approved follow-up, and approved follow-up to CMMS handoff. Report median and range, not only an average. Segment laboratory samples, online alarms, and planned assessments. A faster draft is not a benefit if engineers spend more time correcting it.
3. Decision quality
Use reviewer outcomes as the reference: accepted, edited, rejected, deferred, escalated, or returned for more evidence. Track how often the draft exposed a missing record, linked the right source, distinguished illustrative assumptions from measured values, or surfaced contradictory evidence. Record whether the final decision changed after human review. Do not treat model agreement with a threshold as correctness.
4. Maintenance execution
Measure the percentage of approved packages that reach the CMMS with complete scope, safety and outage assumptions, evidence links, and as-found/as-left fields. Track rework caused by missing context, scope changes, duplicate work, and deferred tasks. CIGRE TB 962 places maintenance in a cycle of planning, execution, recording, and optimization; the pilot should measure that whole cycle.
5. Governance and safety
Track reviewer identity, approval time, permission scope, source access, data-retention status, and exceptions. NIST’s AI RMF provides a useful governance, mapping, measurement, and management frame. If records may include BES Cyber System Information, involve the responsible compliance team and assess NERC CIP-011-3. A pilot metric should never reward bypassing an approval gate.
Illustrative scorecard, not a customer result
Suppose a fictional 16-week pilot covers 12 transformers. Its illustrative scorecard could report: 92% of reviewed cases had an asset-linked source record; median evidence-assembly time fell from a locally measured baseline of 3.5 hours to 2.1 hours; 18 cases were returned for missing context; 7 were escalated after engineer review; and 0 automated OT actions were permitted. These numbers are invented to show the shape of a scorecard, not customer measurements or a performance claim.
The pilot should also report negatives: false or duplicate alerts, reviewer disagreement, stale monitor data, cases with no outcome, and records that could not be safely integrated. The Bureau of Reclamation’s condition-monitoring research emphasizes turning large data volumes into relevant information for operations and maintenance; relevance is a design outcome to test, not a guaranteed benefit.
Commissioning and maintenance workflow
Commission the metric definitions with sample cases. Give reviewers an illustrative case, a known-good historical case, a contradictory case, and a case with missing timestamps. Confirm that the system preserves raw values, labels calculated rates, and routes unresolved cases to a human queue. Test the CMMS handoff in staging and verify that no draft can become an approved order without the utility’s control.
During maintenance, capture as-found and as-left results and link them back to the pilot case. Compare the decision against what was known at the time, not against later hindsight. A DGA alarm remains a review trigger; IEEE C57.104 and IEEE C57.143 inform interpretation and monitoring application, but neither turns a threshold into an automatic trip or maintenance order.
Product boundary
AgenticGrid Pro is a bounded transformer APM workbench that can organize permitted records, calculate configured indicators, show provenance gaps, and draft pilot scorecards for review. ProtectionAI is GridAPM’s separate Windows relay-testing application; it can contribute test records but is not a CBM authority. Neither product promises certification, autonomous operation, guaranteed savings, or replacement of qualified engineers or physical test equipment. Final decisions and OT actions remain engineer-approved.
Use the condition-based maintenance guide, explainable health index article, and pilot page to define the cohort and the review boundary before collecting KPI results.
References
- CIGRE. (2019). Condition assessment of power transformers (Technical Brochure 761). https://www.e-cigre.org/publications/detail/761-condition-assessment-of-power-transformers.html
- CIGRE. (2025). Guide for transformer maintenance (Technical Brochure 962). https://www.e-cigre.org/publications/detail/962-guide-for-transformer-maintenance.html
- IEEE. (2019). IEEE guide for the interpretation of gases generated in mineral oil-immersed transformers (IEEE Std C57.104-2019). https://standards.ieee.org/ieee/C57.104/7476/
- IEEE. (2024). IEEE guide for application of monitoring equipment to liquid-immersed transformers and components (IEEE Std C57.143-2024). https://standards.ieee.org/ieee/C57.143/6961/
- Tabassi, E. (2023). Artificial intelligence risk management framework (AI RMF 1.0) (NIST AI 100-1). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.100-1
- North American Electric Reliability Corporation. (n.d.). CIP-011-3: Cyber Security—Information Protection. Retrieved August 1, 2026, from https://prod.nerc.com/standards/reliability-standards/cip/cip-011-3
- U.S. Bureau of Reclamation. (n.d.). Condition monitoring data acquisition systems. https://www.usbr.gov/research/projects/detail.cfm?id=2879
- Thang, K. F., Aggarwal, R. K., McGrail, A. J., & Esp, D. G. (2003). Analysis of power transformer dissolved gas data using the self-organizing map. IEEE Transactions on Power Delivery, 18(4), 1241–1248. https://doi.org/10.1109/TPWRD.2003.817733
References
- CIGRE TB 761 CIGRE TB 761: Condition assessment of power transformers
- CIGRE TB 962 CIGRE TB 962: Guide for transformer maintenance
- IEEE C57.104 IEEE C57.104-2019: Guide for the Interpretation of Gases Generated in Mineral Oil-Immersed Transformers
- IEEE C57.143 IEEE C57.143-2024: Application of Monitoring Equipment to Liquid-Immersed Transformers
- NIST AI RMF NIST AI RMF 1.0
- NERC CIP-011 NERC CIP-011-3: Cyber Security—Information Protection
- U.S. Bureau of Reclamation condition monitoring data acquisition research
- Thang et al., Analysis of power transformer dissolved gas data using the self-organizing map
Questions engineers ask
What should a transformer CBM pilot measure first?
Start with source completeness, time to assemble evidence, review latency, missing-context detection, reviewer edits or escalations, and approved follow-up. These leading indicators can be measured before any long-term failure or cost claim.
Can a CBM pilot claim guaranteed savings?
No. Savings, outage avoidance, and failure reduction require a defined baseline, comparable cohort, agreed cost model, and enough follow-up data. A pilot should report observed results for its scope, not universal guarantees.
How should a DGA alarm count in pilot metrics?
Count it as a review trigger and measure the quality and timeliness of the review. A DGA threshold should not be counted as an automatic trip or maintenance order.


