Agentic AI & governance

Utility AI Pilot Procurement Checklist: Questions for Security, OT, and Engineering Review

A concrete utility AI pilot procurement checklist covering local-first deployment, controlled egress, role-based review, evidence retention, model logging, acceptance tests, and rollback.

Utility procurement, OT security, and engineering reviewers checking an AI pilot acceptance and rollback checklist
On this page

A utility AI pilot should be procured as a controlled engineering capability, not as a demo with a new interface. The buyer needs evidence that the system can remain inside the approved security boundary, preserve source records, expose uncertainty, support role-based review, and stop cleanly when a test or assumption fails.

The questions below are intended for procurement, CISO, OT security, engineering, reliability, and operations reviewers. They apply to transformer APM, protection evidence, planning support, and other high-consequence workflows. The final decision is always the utility’s: a pilot may prepare evidence, but it does not receive authority merely because a model produces plausible text.

Start with the decision and the boundary

Ask the supplier to state the exact workflow, inputs, outputs, users, excluded uses, and approval points. Require a written answer to these questions:

  • Which engineering decision is supported, and which decisions are out of scope?
  • Does the system read evidence, draft a recommendation, create a review item, or write to an operational system?
  • Can it issue a control, trip, block, setpoint, switching, or maintenance-dispatch command? The acceptable answer for this pilot should be no.
  • Which qualified roles review, edit, approve, reject, or escalate an output?
  • What happens when the source data is missing, stale, contradictory, or outside the model’s evaluation set?

IEEE 7000 is useful here because it connects stakeholder values to traceable system requirements. Convert values such as safety, accountability, privacy, reliability, and contestability into observable requirements and acceptance tests. “Responsible AI” should not remain a slogan in the procurement file.

Architecture and egress questions

NIST SP 800-82 Rev. 3 emphasizes that OT security must account for performance, reliability, and safety constraints. Ask for an architecture diagram that separates the AI application, source repositories, identity provider, integration layer, OT networks, and any vendor-operated service.

Procurement questionEvidence to require
Is deployment local-first?Data-flow diagram, installation boundary, runtime dependencies, and a documented offline or disconnected operating mode where required
What leaves the utility boundary?Field-level egress inventory, destinations, encryption, retention, operator control, and default-deny behavior
Can external model calls be disabled?Configuration, test procedure, and logs showing that the workflow remains usable without unapproved calls
How are identities and roles enforced?Role matrix, least-privilege design, review permissions, separation of duties, and access logs
Does the product connect to OT?Interface inventory, protocol and directionality, write permissions, safety case, and a hard statement that no autonomous control is permitted

NIST CSF 2.0 can structure the cybersecurity conversation across governance, identification, protection, detection, response, and recovery. NERC CIP obligations may apply depending on the entity, system, and jurisdiction; the supplier should support the utility’s compliance analysis rather than claim certification on the buyer’s behalf. IEEE 1686 is a useful reference when the pilot touches IED configuration, firmware, access, or data retrieval.

Evidence retention and model logging

Require every material output to retain the input record IDs, source timestamps, prompt or workflow version where applicable, model and retrieval version, user identity, role, approval state, edits, and final disposition. The record should distinguish source evidence, generated text, engineer edits, and approved action.

Ask whether the buyer can export the complete evidence pack in a usable format. Ask how long records are retained, how legal hold or investigation requests are handled, how deleted or superseded sources are represented, and whether a reviewer can reconstruct the output after a model or prompt update. A new model version should not silently change the meaning of old approvals.

Harvard Data Science Review research on human-centered AI transparency is a useful warning: transparency should support appropriate human understanding and action, not merely expose technical detail. An explanation that cannot help an engineer inspect, challenge, or reverse an output is not a sufficient control.

Acceptance tests that matter

Write acceptance tests against real workflow evidence, with synthetic or sanitized data when necessary. Include at least:

  1. Identity and provenance: the system maps records to the correct asset, preserves timestamps and units, and shows the source for every material statement.
  2. Missing or conflicting evidence: the system flags the gap, avoids inventing a conclusion, and routes the case to a named reviewer.
  3. Role-based review: an unapproved user cannot approve a recommendation, change an OT boundary, or bypass a required hold point.
  4. No autonomous control: the pilot cannot issue protection, switching, setpoint, or maintenance-dispatch commands, including through an integration side effect.
  5. Model and version logging: the output records the application, model, prompt or workflow, retrieval corpus, and configuration versions used.
  6. Evidence retention: an exported package contains source references, generated content, reviewer edits, approvals, and final disposition.
  7. Rollback: the buyer can disable the pilot, revoke credentials, restore the prior workflow, and preserve the records without data loss.
  8. Operational failure: timeout, unavailable model, bad file, malformed record, and permission failure produce a safe, visible error rather than a partial action.

Do not accept a benchmark that measures only answer fluency. For protection, planning, and maintenance, acceptance should test traceability, reviewer behavior, false confidence, boundary enforcement, and recovery.

Product-specific scope

ProtectionAI is Windows desktop protective-relay testing software with an agentic AI copilot. A procurement review can examine its settings, test-plan, COMTRADE, SCL/RIO/XRIO, report, and offline-job-pack workflows within documented scope. It does not control physical test sets, provide on-network GOOSE or Sampled Values, or replace qualified protection engineers or physical test equipment.

AgenticGrid Pro is the power-transformer APM workbench. A pilot can evaluate its source-linked condition, maintenance, event, and review workflows; it should not be sold as autonomous operation, final diagnosis, work-order authority, or a substitute for field inspection and engineering judgment.

The GridAPM procurement page, security model, data handling, trust model, and pilot evaluation provide internal review context. The OT AI risk register and AI permission model can help convert the checklist into a pilot control register.

References

References

  1. NIST SP 800-82 Rev. 3 — Guide to Operational Technology Security
  2. NIST AI RMF NIST AI Risk Management Framework
  3. NIST Cybersecurity Framework 2.0
  4. IEEE 7000 IEEE 7000-2021 — Model Process for Addressing Ethical Concerns During System Design
  5. IEEE 1686 IEEE 1686-2022 — Intelligent Electronic Devices Cybersecurity Capabilities
  6. CIGRE Technical Brochure 946 — AI/ML in power network operation and control
  7. IEEE PES-TR112 — AI/ML in Power System Protection and Control
  8. Harvard Data Science Review — AI Transparency in the Age of LLMs
  9. NERC — Critical Infrastructure Protection Reliability Standards

Questions engineers ask

What is the first question in a utility AI pilot procurement?

Ask what decision the pilot supports, what it explicitly cannot do, which approved users may review it, and what evidence proves usefulness. A general promise of automation is not a bounded utility use case.

Should a utility require local-first deployment?

Local-first should be evaluated whenever operational, security, privacy, or data-residency requirements make external processing unacceptable. The buyer should also require a documented egress design, approved destinations, logging, and a way to operate without uncontrolled cloud access.

What should happen if the pilot fails an acceptance test?

The pilot should fail closed for the affected workflow, preserve the test evidence, prevent unapproved promotion, and support rollback to the prior approved process. A failed test should not be hidden by changing the success definition after the fact.

Filed under

AI procurementOT securityUtility cybersecurityPilot evaluationModel governanceEvidence retentionHuman-in-the-loop

Discuss this with our engineers

Share your fleet profile and diagnostic workflow. GridAPM will propose a focused pilot evaluation path.

Type to search research, platform pages, and tools.