Abstract

AI can help a utility accomplish more when it improves the complete path from information to an accepted outcome. This paper examines that proposition across customer service, planning, asset work, and operational support. Drawing on Salesforce’s utilities guide and public DOE and NIST guidance, we propose a framework for selecting workflows, testing their performance, and measuring their contribution to the business. The framework is a research proposal; it does not report a customer trial or demonstrated savings.

CHAPTER 01

An application is the beginning of a business case.

Salesforce’s AI in Utilities guide describes opportunities in customer service, demand forecasting, maintenance, outage response, energy procurement, and renewable integration. It also identifies privacy, integration, and responsible decision-making as implementation challenges. These are useful directions for investigation. The guide does not establish what a particular utility will save, how much human review it will need, or whether a specific deployment will meet its operating requirements.

The U.S. Department of Energy’s 2024 AI for Energy report similarly identifies opportunities in planning, permitting, operations and reliability, and resilience. Its examples include faster grid studies and support for permitting review. This expands the question beyond a single assistant: where could AI increase the organization’s ability to deliver?

Our research question is practical: under what conditions does an AI-assisted workflow produce more acceptable work with the same resources, while preserving the quality and authority the work requires?

Corvana’s proposed answer begins with a bounded workflow. “Customer service” is a department. “Prepare a billing explanation from an approved account record and tariff, then route exceptions to a specialist” is work that can be tested. That distinction makes the ambition measurable.

CHAPTER 02

Select work with a visible outcome.

The application areas below draw on Salesforce’s guide; the workflow boundaries and evaluation measures are Corvana’s proposals. They are candidates for testing, not descriptions of deployed Corvana products.

From an application area to a testable workflow
Application areaA bounded starting pointEvidence to collect
Customer servicePrepare a sourced billing explanation for an agent to review.Resolution quality, handling time, repeat contact, and escalation.
Demand forecastingPrepare a forecast comparison and flag uncertain periods for planners.Error by season and peak period; planner effort and overrides.
Asset maintenanceAssemble inspection history and a draft work package.Missing evidence, planner corrections, and time to an accepted package.
Outage managementSummarize approved incident records for a restoration briefing.Factual support, information freshness, preparation time, and omissions.
Renewables and storageCompare generation scenarios and document their assumptions.Forecast uncertainty, scenario coverage, and analyst review effort.
Procurement and pricingPrepare a comparison of approved offers against stated constraints.Calculation accuracy, source coverage, and analyst corrections.

These workflows may require different techniques. A numerical forecast, an equipment anomaly detector, and a document-drafting assistant should not share a generic “AI accuracy” score. Each needs a baseline and an outcome appropriate to its job.

We propose prioritizing candidates with repeatable demand, accessible source records, an identifiable owner, and an outcome that qualified people can grade. A costly bottleneck with unusable records may first need data work. A polished demonstration with little recurring demand may have limited economic value. Selection should account for both.

FIELD NOTES / INSIDE THE SANDBOX

One case. Every consequence accounted for.

Follow a case from its source records to the decision. Each example includes the point where human judgment changes the outcome.

Fictional walkthroughs for illustration. These are proposed tests, not customer results.

SCENARIO 02 / MISSING-INFORMATION DETECTION

Does the system know when the record is incomplete?

A planner replays a completed maintenance package. The agent sees the approved asset history and procedures available at the time, with one inspection attachment deliberately omitted.

STEP 01

An incomplete asset record

Use only the information available at the time. Freeze the sources, withhold the final answer, and preserve an untouched set of cases for evaluation.

STEP 02

Assemble the package; expose the gap

The agent produces a draft checklist and an evidence index. It marks the missing inspection as unresolved instead of inventing a reading or a maintenance recommendation.

STEP 03

Stopping can be the correct result.

The planner checks asset identity and procedure revision, confirms that the missing evidence was detected, and records the work needed to complete the package manually.

STEP 04

Pass this case; continue evaluation

The missing-record control worked in this fictional case. The full case set must still meet quality and effort thresholds before any rollout decision.

What we measure

Missing-record detection · package completeness · planner time

The boundary

The sandbox cannot dispatch a crew or operate equipment.

SCENARIO 01 / SOURCE-BASED DRAFTING

Can a faster response survive the evidence check?

A regulatory team replays a closed data request about vegetation-management spending. The final filed response stays hidden from the agent.

STEP 01

A closed data request

Use only the information available at the time. Freeze the sources, withhold the final answer, and preserve an untouched set of cases for evaluation.

STEP 02

Prepare a response with citations

The agent assembles the cost explanation and links each claim to a ledger row or work order. It flags a mismatch in reporting periods.

STEP 03

A convincing draft can still fail.

The reviewer finds that one cost total includes work outside the requested period. The draft goes back for correction. That review time stays in the result.

STEP 04

Revise & retest

Correct the date filter, then rerun on a held-out case set. A lower drafting time cannot compensate for an unsupported total.

What we measure

Citation support · period accuracy · reviewer effort

The boundary

The sandbox cannot file a response or change the ledger.

SCENARIO 03 / POLICY & EXCEPTION HANDLING

Can it explain the bill without inventing a policy?

A customer-operations team replays a closed billing query across a tariff change. The account record is synthetic, and the final agent response is withheld.

STEP 01

A bill across two tariff periods

Use only the information available at the time. Freeze the sources, withhold the final answer, and preserve an untouched set of cases for evaluation.

STEP 02

Explain the charges against the record

The agent drafts a line-by-line explanation, cites each effective rate, and identifies a disputed adjustment that needs a specialist.

STEP 03

Check the exception, not just the arithmetic.

A reviewer confirms the rates and calculation, then checks that the disputed adjustment was escalated. A plausible but inapplicable tariff would fail the case.

STEP 04

Proceed to a limited shadow test

If the full evaluation clears its agreed bar, compare drafts with the existing team in a limited shadow run. A person still owns every customer-facing response.

What we measure

Rate accuracy · appropriate escalation · total handling time

The boundary

The sandbox cannot message a customer or alter an account.

CHAPTER 03

Test the whole path to an accepted result.

The proposed evaluation unit is the completed case, including human review and any fallback. If an assistant drafts a package quickly but shifts work to an engineer who must reconstruct its evidence, a model-level speed improvement has not established a workflow improvement.

For each candidate, Corvana proposes an evaluation with five elements:

  1. Define the operating boundary. Name the task owner, permitted records, required output, and actions the system may take. For an initial document workflow, use an isolated environment with approved records and no live transaction authority. State what remains a human decision.
  2. Build a comparison set. Separate development cases from cases reserved for evaluation. Reconstruct the information available when each case was handled; hide later outcomes. Include ordinary work, incomplete records, conflicting sources, and cases previously escalated. For forecasts, preserve the order of time and test unusual operating periods separately.
  3. Agree on acceptance before running. Have domain reviewers define which errors require rejection, correction, or escalation. For a billing explanation, check account facts and the applicable tariff. For a maintenance package, check asset identity, document completeness, and unsupported recommendations. A persuasive answer can still fail.
  4. Measure completion and effort. Compare equivalent cases with and without AI assistance. Record preparation, verification, corrections, escalations, failed attempts, and manual fallback. Separate hands-on time from waiting time. Report results by case type, including the cases the workflow could not finish.
  5. Report uncertainty and a decision. Retain the input, model and tool versions, reviewer scores, and reasons for failure. Examine variation across repeated runs and reviewers. A small sample without observed serious errors is limited evidence. End with a recommendation to proceed within scope, revise and retest, or stop.

NIST’s AI Risk Management Framework Core supports documenting context, benchmarks, oversight, evaluation methods, and continuing monitoring. It organizes risk management through Govern, Map, Measure, and Manage. We use it as guidance for structuring evidence, not as proof that an individual workflow is safe or effective.

THE EVALUATION BOUNDARYNO LIVE ACTIONS
01

Freeze the record

Approved inputs.
Final answers withheld.

02

Replay the work

Bounded tools.
Every attempt logged.

03

Review the result

Accepted outcomes.
Total effort measured.

FIG. 01 The sandbox reproduces the work while keeping consequential actions outside the evaluation.

CHAPTER 04

More capacity and lower spending are different outcomes.

We propose keeping three ledgers: work completed, service and decision quality, and cash impact. They answer different questions. A team may clear a backlog without reducing payroll. Faster preparation may improve turnaround while a later approval remains the limiting step. Both can matter, but neither should automatically be reported as an expense reduction.

Proposed measurement model

Annual hours released
Eligible cases × actual adoption × net hands-on hours saved per case

Net recurring cash benefit
Verified reduction in spending − incremental recurring operating cost

Net hours saved must include review, rework, and fallback. Recurring costs include models, tools, monitoring, support, and retesting. One-time evaluation, integration, and training costs belong in the investment case as well. Keep each cost in one place to avoid double counting.

Before rollout, finance and the workflow owner should identify how capacity will be used: completing deferred planning work, improving response times, reducing overtime, or changing a specific contractor commitment. Only the spending changes belong in cash savings. Forecasts of adoption and benefit should be tested against actual use after deployment.

For a worked financial example, see The economic case for an AI sandbox. The same discipline applies here: the economic case depends on the entire process, not the price of a model call.

CHAPTER 05

Match the next step to the consequence.

A useful sandbox result earns a deployment decision. It does not itself grant new permissions. We propose separating three kinds of authority when planning the next test.

Prepare information

The workflow retrieves, organizes, and drafts within a defined information boundary. Test source access, freshness, omissions, and the reviewer’s ability to trace a conclusion. Keep a clear route to reject or correct the result.

Support a decision

The workflow compares options or recommends a course of action. Test it against domain-specific criteria and counterexamples. Make uncertainty visible, record the reviewer’s disposition, and measure whether recommendations improve the decision process rather than simply getting accepted.

Take an operational action

Changing equipment settings, executing a transaction, or issuing a consequential instruction requires additional assurance specific to the system involved. A document replay cannot establish readiness for live grid control. The relevant engineering and operational owners must define the tests, safeguards, and authority appropriate to that action.

For a workflow that proceeds, start with a limited group and a defined review period. Compare actual quality, adoption, cost, and exception handling with the evaluation. Set conditions for pausing or reverting the workflow, and reassess material changes to models, tools, permissions, or source data.

CHAPTER 06

Build the company’s ability to deliver.

Our central proposition is that AI creates durable operating value when a company can repeatedly convert domain knowledge into work it can trust and use. The first successful workflow can leave behind reusable evaluation cases, clear source ownership, review criteria, and a measured baseline.

That foundation can make the next evaluation easier, but its results do not transfer automatically. A billing workflow does not validate a maintenance recommendation. Expansion still requires new evidence for the new context.

Start with work you can measure. Build toward a company that can accomplish more.

For Corvana, that is the ambition: connect decades of operating expertise to additional capacity, with evaluation making each step accountable.

Sources, method & limitations

This paper is a selective literature synthesis and an original proposed evaluation framework. It uses Salesforce’s guide to frame application areas, DOE’s report to provide energy-sector context, and NIST’s framework to inform evidence and governance. It is not a systematic review, a peer-reviewed empirical study, or a report of Corvana customer results. No utility dataset was analyzed and no performance uplift was measured. The proposed criteria and economics require validation in the intended operating setting.

  1. Salesforce. AI in Utilities. Industry guide; the starting point for this paper. Accessed September 10, 2026.
  2. U.S. Department of Energy. AI for Energy: Opportunities for a Modern Grid and Clean Energy Economy. April 29, 2024. Accessed September 10, 2026.
  3. National Institute of Standards and Technology. AI RMF 1.0: Core. 2023. Accessed September 10, 2026.

The workflow proposals and interpretations are Corvana’s. The cited organizations have not endorsed this paper.