Insights · Case Studies

Automotive · Document intelligence

Finding the revisions that actually change the repair

Procedure Change Detection

25,000+ changes, each traced to its step.

Fewer than 3% of pairs need a domain expert.

Client type
Confidential · Service-Data Provider
Stack
LLMs + Databricks + MLflow

Surface every revision that actually changes the repair, with evidence down to the step, and let the rest of the catalog pass through untouched.

Every model year, an automotive service-content catalog is re-issued and thousands of repair procedures are revised. Most edits are cosmetic, but hidden among them are the added steps, revised torque values, and merged jobs that move labor time. Re-estimating labor is expensive, expert-dependent work. It should only ever touch those revisions, not the entire catalog.

The Challenge

A catalog that changes every year, mostly in ways that don't matter

Noise Dominates

Reformatting, renumbering, and rewording make up the vast majority of edits, and none of them change what a technician actually does.

Signal Is Rare and Scattered

Only a small fraction of revisions actually change the repair and its labor time, and they can hide anywhere in the text or the figures.

Manual Review Doesn't Scale

Re-reading two versions of thousands of procedures is slow, and different reviewers disagree about what counts as a real change.

The Approach

Read both years, name every change, prove it

Two kinds of evidence stand behind every verdict: AI models make the judgment calls, and model-free code produces scores anyone can re-run.

Every difference is cited to the exact step or figure it came from, so any verdict can be opened and checked down to the line.

System Architecture

How the pipeline works

  1. Two model years

    290K+ configurations in scope

    Every procedure in both the outgoing and incoming model year: full text, figures, and referenced procedures.

  2. Pair and filter

    290K+ → ~3,400 changed

    A content hash drops the byte-identical pairs, so only genuine edits move forward.

  3. Classify, quantify, trace

    25,000+ changes, 100% traced

    A primary model names and sizes every difference, deterministic scores measure it, and each change is cited to its exact step or figure.

  4. Route

    ~70% of uncertain auto-resolved

    Clear pairs settle by rule, a judge model resolves most uncertain ones, and only the truly ambiguous reach a domain expert.

  5. Auditable change set

    <3% need a domain expert

    A filterable record of what changed and how much, feeding labor re-estimation with evidence attached.

By the Numbers

One model-year refresh, measured end to end

290K+

vehicle configurations matched across both model years

25,000+

individual changes named and traced

83%

of pairs where both models agree on the category

<3%

of pairs need a domain expert

Figures are from a single full-population run and are rounded; they describe the pipeline's operation, not a client's business results.

Deployment & Iteration

Built to run across the whole catalog

A refresh this size is as much an operations problem as a modeling one, so the pipeline treats it that way: one governed, repeatable job, with its expert-routing threshold calibrated rather than guessed.

  1. Full-catalog scale

    Executed across the entire changed-procedure population on Databricks, in an isolated, fully logged run.

  2. Resumable and observable

    Chunked, restartable passes with per-run metrics, so a long job survives interruptions and its progress is visible throughout.

  3. Calibrated with experts

    A stratified sample is reviewed by domain experts to tune where the line for manual validation should sit.

Outcome

Every real change, found and accounted for

~3,400

from 290K+ matched pairs

Noise Never Reaches a Reviewer

The content-hash filter collapses hundreds of thousands of matched configurations down to a few thousand actually changed procedures. The cosmetic churn never lands on a human desk.

25,000+

changes, 100% cited

Every Change Has an Address

Each difference is named, sized, and tied to the exact step or figure it came from. Nothing gets surfaced without its evidence.

<3%

of pairs need an expert

Experts See Only the Hard Cases

The judge model settles about 70% of the uncertain pairs automatically; only the truly ambiguous few reach a domain expert.

“The result is an auditable record of what changed and how much, from a full model-year refresh down to a single step, so labor re-estimation targets only the revisions that matter.”

Technical Strength

Deep expertise across the document-intelligence stack

Document Intelligence & Structured Diffing

Figure labelling, step numbering, and content-hash filtering that make two years of procedures comparable and citable.

Multi-Model LLM Orchestration

A primary classifier and an independent judge, with rules deciding what settles automatically and what escalates.

Deterministic Evaluation & Scoring

Reproducible, model-free scores that sit beside every model judgment.

Databricks-Scale Pipeline Engineering

Chunked, restartable, fully logged runs across the entire catalog.

Work with us

Have a problem worth publishing?

If it's hard enough to end up in this record, it's the kind of work we want. Tell us what to build, and we'll show ours.