Custom AI Models & Validation

Find out whether AI
works for your workflow.

Compare vendors, assess an existing assistant or pilot, or verify an update on your own work. We evaluate your system, another vendor’s, or a Corvana workflow and report quality, failures, and reviewer effort.

Discuss an evaluation
Explore the approach

What your team receives

Know where AI works.
See where it breaks.

Reusable test cases, scoring criteria, graded outputs, and a findings report. Evidence for your next decision, with tests your team can run again.

01 / Compare vendors

Reusable tasks.
Your own work.

Receive a task pack designed with your practitioners: source records, expected outputs, and acceptance criteria to compare systems on the same work.

02 / Assess an assistant

Graded outputs.
Clear evidence.

Assess an existing assistant or pilot with agreed scoring criteria, graded outputs, and failure evidence. See which results need correction or human review.

03 / Verify an update

Findings report.
Tests to re-run.

Receive a report covering quality, critical failures, reviewer effort, and cost. Re-run the tests after model, data, or workflow changes to see what changed.

How we evaluate your AI workflow

Real utility work.
A standard you can test.

Operational knowledge supplies the records, procedures, and relationships. Evaluation tests whether AI uses them correctly. Each service can be scoped independently.

01 / Define the work

Start with work your
utility already knows.

Start with the decision you need to make: choose a vendor, assess an assistant, or verify an update. With your practitioners, we scope one workflow, establish its current performance, and agree on the evidence and acceptance criteria. The contractor-readiness example below shows what we test.

  • One bounded workflow
  • Approved source material
  • Expert-defined acceptance criteria
Contractor readiness · CP–104Illustrative example · Synthetic records
Contractor readiness / Package CP–104

Is the evidence valid for mobilization?

Review CP–104 using its certificate, P6 milestone MOB–104, and contractor readiness procedure §4. Determine whether the evidence supports planned mobilization, flag unresolved issues, and prepare a cited review for the named capital delivery approver.

Certificate expiry 24 SeptemberMobilization 26 SeptemberProcedure §4 Cover at mobilization
A demonstration of evaluation design, not a model performance result.

Research context: EPRI distinguishes power-sector knowledge tests from future workflow evaluations. PNNL identifies handoffs, maintenance records, and outage documentation as candidate utility applications. EPRI · WattWorks ↗ · PNNL · Utility applications ↗

02 / Build the task set

Give every task
its working context.

Package the brief, source documents, permitted tools, and expected work product into a reproducible case. Include everyday work, conflicting records, outdated revisions, and missing information. Keep the final test set separate from development and training data.

  • Authorized or synthetic records
  • Typical and difficult cases
  • Private, versioned holdout set
Inside a benchmark caseIllustrative structure
Task brief

Review the record.
Show your evidence.

01Source bundle

Files, revisions, provenance

02Work environment

Tools, access, run limits

03Grading criteria

Expected output, checks

Development casesHeld-out evaluation cases

Method reference: Legora BAR ↗ evaluates complete tasks with source files and expert-written rubrics. We apply that task-based principle to utility workflows; the examples here are our own proposed designs.

03 / Grade the outcomes

Make every failure
specific and actionable.

Your practitioners help define and weight the rubric. Code checks exact fields, calculations, units, and references where possible. Expert review handles judgment; model graders are checked against human decisions before use.

  • Evidence and numerical accuracy
  • Completeness and uncertainty
  • Access and action boundaries
What could we evaluate?Illustrative task designs
01 / Regulatory work

Draft a response. Account for every claim.

Prepare a response to a data request using the supplied filing, testimony, and supporting schedules. Flag anything the evidence cannot resolve.

Source bundle
Approved filing · Testimony · Schedules
Work product
A cited draft and a list of unresolved questions.
What we check
Citations support claims. Reporting periods agree. Totals and units reconcile.

Examples of benchmark design, not client projects or performance results.

Critical failures stay visible. A strong average quality score cannot cancel out an access violation or an unauthorized action.

Method references: Anthropic · Agent evaluations ↗ · EPRI · Use-case risk and autonomy ↗

04 / Compare systems

See what changed.
Understand what improved.

Compare models under the same tasks, rubrics, tools, and budget. When testing a new retrieval setup or agent workflow, record the full system configuration. Repeat runs to measure variability, and inspect performance by task type and failure severity.

  • Comparable run conditions
  • Quality, latency, and cost
  • Critical failures reported separately
One benchmark. A clearer comparison.
Evaluation reportReport structure · No sample scores
Quality
Correctness · Evidence · Completeness
Reliability
Variation across repeated runs
Boundaries
Critical failures and their traces
Efficiency
Time to completion · Cost per task

Method references: EPRI · Repeatability testing ↗ · Anthropic · Evaluating agent systems ↗

05 / Keep improving

Build once.
Learn from every run.

Receive an evaluation report covering acceptance criteria, failure modes, reviewer effort, and cost per task, alongside the task pack, grading specification, and run configuration. Use the findings to decide what to adopt, improve, or test further. Re-run the benchmark after model, data, or workflow changes, keeping a separate holdout for final validation.

  • Reusable task and rubric pack
  • Documented run setup
  • Evidence for the next improvement
A repeatable improvement cycle
  1. 01Evaluate the workflow
  2. 02Review the failure evidence
  3. 03Improve. Version. Re-run.
Keep the standard stable. Make the system better.

Method reference: NIST AI RMF Playbook · Measure ↗

Results apply to the tasks and conditions tested. Operational control needs additional domain validation, simulation, and deployment review; a language-model benchmark does not establish grid safety.

Your work. Your standard.

Know what your AI
can actually do.

Bring a vendor choice, an existing assistant, or a planned update.
We’ll define the tests that support your next decision.

Discuss an evaluation