Utility AI Pilot Procurement Checklist: Questions for Security, OT, and Engineering Review
A concrete utility AI pilot procurement checklist covering local-first deployment, controlled egress, role-based review, evidence retention, model logging, acceptance tests, and rollback.

On this page
A utility AI pilot should be procured as a controlled engineering capability, not as a demo with a new interface. The buyer needs evidence that the system can remain inside the approved security boundary, preserve source records, expose uncertainty, support role-based review, and stop cleanly when a test or assumption fails.
The questions below are intended for procurement, CISO, OT security, engineering, reliability, and operations reviewers. They apply to transformer APM, protection evidence, planning support, and other high-consequence workflows. The final decision is always the utility’s: a pilot may prepare evidence, but it does not receive authority merely because a model produces plausible text.
Start with the decision and the boundary
Ask the supplier to state the exact workflow, inputs, outputs, users, excluded uses, and approval points. Require a written answer to these questions:
- Which engineering decision is supported, and which decisions are out of scope?
- Does the system read evidence, draft a recommendation, create a review item, or write to an operational system?
- Can it issue a control, trip, block, setpoint, switching, or maintenance-dispatch command? The acceptable answer for this pilot should be no.
- Which qualified roles review, edit, approve, reject, or escalate an output?
- What happens when the source data is missing, stale, contradictory, or outside the model’s evaluation set?
IEEE 7000 is useful here because it connects stakeholder values to traceable system requirements. Convert values such as safety, accountability, privacy, reliability, and contestability into observable requirements and acceptance tests. “Responsible AI” should not remain a slogan in the procurement file.
Architecture and egress questions
NIST SP 800-82 Rev. 3 emphasizes that OT security must account for performance, reliability, and safety constraints. Ask for an architecture diagram that separates the AI application, source repositories, identity provider, integration layer, OT networks, and any vendor-operated service.
| Procurement question | Evidence to require |
|---|---|
| Is deployment local-first? | Data-flow diagram, installation boundary, runtime dependencies, and a documented offline or disconnected operating mode where required |
| What leaves the utility boundary? | Field-level egress inventory, destinations, encryption, retention, operator control, and default-deny behavior |
| Can external model calls be disabled? | Configuration, test procedure, and logs showing that the workflow remains usable without unapproved calls |
| How are identities and roles enforced? | Role matrix, least-privilege design, review permissions, separation of duties, and access logs |
| Does the product connect to OT? | Interface inventory, protocol and directionality, write permissions, safety case, and a hard statement that no autonomous control is permitted |
NIST CSF 2.0 can structure the cybersecurity conversation across governance, identification, protection, detection, response, and recovery. NERC CIP obligations may apply depending on the entity, system, and jurisdiction; the supplier should support the utility’s compliance analysis rather than claim certification on the buyer’s behalf. IEEE 1686 is a useful reference when the pilot touches IED configuration, firmware, access, or data retrieval.
Evidence retention and model logging
Require every material output to retain the input record IDs, source timestamps, prompt or workflow version where applicable, model and retrieval version, user identity, role, approval state, edits, and final disposition. The record should distinguish source evidence, generated text, engineer edits, and approved action.
Ask whether the buyer can export the complete evidence pack in a usable format. Ask how long records are retained, how legal hold or investigation requests are handled, how deleted or superseded sources are represented, and whether a reviewer can reconstruct the output after a model or prompt update. A new model version should not silently change the meaning of old approvals.
Harvard Data Science Review research on human-centered AI transparency is a useful warning: transparency should support appropriate human understanding and action, not merely expose technical detail. An explanation that cannot help an engineer inspect, challenge, or reverse an output is not a sufficient control.
Acceptance tests that matter
Write acceptance tests against real workflow evidence, with synthetic or sanitized data when necessary. Include at least:
- Identity and provenance: the system maps records to the correct asset, preserves timestamps and units, and shows the source for every material statement.
- Missing or conflicting evidence: the system flags the gap, avoids inventing a conclusion, and routes the case to a named reviewer.
- Role-based review: an unapproved user cannot approve a recommendation, change an OT boundary, or bypass a required hold point.
- No autonomous control: the pilot cannot issue protection, switching, setpoint, or maintenance-dispatch commands, including through an integration side effect.
- Model and version logging: the output records the application, model, prompt or workflow, retrieval corpus, and configuration versions used.
- Evidence retention: an exported package contains source references, generated content, reviewer edits, approvals, and final disposition.
- Rollback: the buyer can disable the pilot, revoke credentials, restore the prior workflow, and preserve the records without data loss.
- Operational failure: timeout, unavailable model, bad file, malformed record, and permission failure produce a safe, visible error rather than a partial action.
Do not accept a benchmark that measures only answer fluency. For protection, planning, and maintenance, acceptance should test traceability, reviewer behavior, false confidence, boundary enforcement, and recovery.
Product-specific scope
ProtectionAI is Windows desktop protective-relay testing software with an agentic AI copilot. A procurement review can examine its settings, test-plan, COMTRADE, SCL/RIO/XRIO, report, and offline-job-pack workflows within documented scope. It does not control physical test sets, provide on-network GOOSE or Sampled Values, or replace qualified protection engineers or physical test equipment.
AgenticGrid Pro is the power-transformer APM workbench. A pilot can evaluate its source-linked condition, maintenance, event, and review workflows; it should not be sold as autonomous operation, final diagnosis, work-order authority, or a substitute for field inspection and engineering judgment.
The GridAPM procurement page, security model, data handling, trust model, and pilot evaluation provide internal review context. The OT AI risk register and AI permission model can help convert the checklist into a pilot control register.
References
- CIGRE. (2024). The impact of the growing use of machine learning/artificial intelligence in the operation and control of power networks from an operational perspective (Technical Brochure 946). https://www.e-cigre.org/publications/detail/946-the-impact-of-the-growing-use-of-machine-learningartificial-intelligence-in-the-operation-and-control-of-power-networks-from-an-operational-perspective.html
- IEEE Power System Relaying and Control Committee. (2023). Practical applications of artificial intelligence and machine learning in power system protection and control (PES-TR112). https://www.pes-psrc.org/kb/report/117.pdf
- Institute of Electrical and Electronics Engineers. (2022). IEEE standard for intelligent electronic devices cybersecurity capabilities (IEEE Std 1686-2022). https://standards.ieee.org/ieee/1686/7207/
- Institute of Electrical and Electronics Engineers. (2021). IEEE standard model process for addressing ethical concerns during system design (IEEE Std 7000-2021). https://standards.ieee.org/ieee/7000/6781/
- Liao, Q. V., & Vaughan, J. W. (2024). AI transparency in the age of LLMs: A human-centered research roadmap. Harvard Data Science Review, Special Issue 5. https://hdsr.mitpress.mit.edu/pub/aelql9qy/release/2
- National Institute of Standards and Technology. (2023). Guide to operational technology (OT) security (NIST Special Publication 800-82 Rev. 3). https://csrc.nist.gov/pubs/sp/800/82/r3/final
- National Institute of Standards and Technology. (2023). Artificial intelligence risk management framework (AI RMF 1.0). https://www.nist.gov/itl/ai-risk-management-framework
- National Institute of Standards and Technology. (2024). The NIST Cybersecurity Framework (CSF) 2.0 (NIST CSWP 29). https://www.nist.gov/publications/nist-cybersecurity-framework-csf-20
- North American Electric Reliability Corporation. (n.d.). CIP: Critical infrastructure protection reliability standards. https://www.nerc.com/standards/reliability-standards/cip
References
- NIST SP 800-82 Rev. 3 — Guide to Operational Technology Security
- NIST AI RMF NIST AI Risk Management Framework
- NIST Cybersecurity Framework 2.0
- IEEE 7000 IEEE 7000-2021 — Model Process for Addressing Ethical Concerns During System Design
- IEEE 1686 IEEE 1686-2022 — Intelligent Electronic Devices Cybersecurity Capabilities
- CIGRE Technical Brochure 946 — AI/ML in power network operation and control
- IEEE PES-TR112 — AI/ML in Power System Protection and Control
- Harvard Data Science Review — AI Transparency in the Age of LLMs
- NERC — Critical Infrastructure Protection Reliability Standards
Questions engineers ask
What is the first question in a utility AI pilot procurement?
Ask what decision the pilot supports, what it explicitly cannot do, which approved users may review it, and what evidence proves usefulness. A general promise of automation is not a bounded utility use case.
Should a utility require local-first deployment?
Local-first should be evaluated whenever operational, security, privacy, or data-residency requirements make external processing unacceptable. The buyer should also require a documented egress design, approved destinations, logging, and a way to operate without uncontrolled cloud access.
What should happen if the pilot fails an acceptance test?
The pilot should fail closed for the affected workflow, preserve the test evidence, prevent unapproved promotion, and support rollback to the prior approved process. A failed test should not be hidden by changing the success definition after the fact.


