Evidence-Gated Engine Layer

AI Transformation Assessment Engine Thinking Flow

Turning untrusted organizational evidence into a verified AI Transformation diagnosis and a confidence-bounded path to value creation. One governed, evidence-gated agentic workflow. Five parallel forensic domains. Separate maturity and anti-pattern streams. Deterministic validation, tactic permissioning, independent fact-checking, and visible GO / WARN / BLOCK boundaries.
Interactive Version
25 Maturity + 25 Anti-Pattern Criteria
150 Evidence Questions
10 Evidence Categories
How to read this visual: the default layer shows the complete control architecture. Each clickable card then reveals the key question, reasoning logic, grounding basis, reliability / hallucination control, and an illustrative thinking sample. The purpose is to show how the Engine decomposes expert AI Transformation reasoning without exposing prompts.
AI reasoning Deterministic control Gate / validation Knowledge / support lane Publication boundary

1. Source intake, parsing & safety boundary

The Engine first converts heterogeneous source material into a safe and attributable assessment pack. It does not begin by asking an AI model whether the organization is ready.

Evidence boundary first
Critical trust idea: customer documents are untrusted evidence. Reference Knowledge Base material can guide rubric interpretation and roadmap mechanisms, but it cannot become proof of the assessed organization’s present state.
PDF
Open logic
Input preparation

Multi-format local parsing

Normalizes PDFs, HTML, tables, JSON, screenshots, and other artifacts before model analysis.

Local extractionPage markersTable rows
Key question
What usable text, visual evidence, table structure, and source location can be extracted from every file?
Reasoning logic
Create machine-readable source material while preserving page, row, image, and document boundaries for later verification.
Grounding basis
Local PDF extraction, sanitized HTML, table profiling, JSON normalization, image inputs, and parse-quality metadata.
Reliability control
The model is not the primary parser. Sparse or ambiguous extraction remains a visible limitation instead of being silently treated as complete evidence.
Illustrative thinking sample“An AI governance clause on page 14 and a platform diagram on page 27 remain separate evidence objects; they cannot become one blended impression.”
DLP
Open logic
Safety control

DLP & sensitive-data review

Scans the distributed source registry for secrets, identifiers, and other unsafe content before deeper analysis.

SecretsDistributed samplingRedaction
Key question
Can the material enter the pipeline without exposing credentials, personal data, or sensitive organizational identifiers?
Reasoning logic
Run deterministic checks and a bounded pre-flight review across the registry, then block or redact high-risk material before model calls.
Grounding basis
DLP pattern hits, registry chunk sampling, image-review limits, safety warnings, and privacy sanitation rules.
Reliability control
Source facts can support internal assessment while names and unnecessary identifying details are removed from report-facing prose.
Illustrative thinking sample“Use the presence of unclear data-access controls as evidence; do not expose named individuals, credentials, or unnecessary entity labels.”
PQ
Open logic
Input quality

Parse quality & visual evidence

Identifies sparse pages, extraction gaps, and visually rich sources that require image-aware interpretation.

Sparse-page warningVisual sourceUsability status
Key question
Does the extracted material adequately represent what a human reviewer can actually see?
Reasoning logic
Carry selected visual pages and screenshots alongside text so architecture diagrams, service maps, dashboards, and operating models are not lost.
Grounding basis
Text coverage, page counts, image normalization, visual-page references, and source-usability classification.
Reliability control
A visually rich but text-poor PDF is labeled visual-only or mixed—not automatically discarded or described as containing no assessable evidence.
Illustrative thinking sample“A service blueprint may prove ownership and handoffs even when its labels appear only inside a rendered diagram.”
KB
Open logic
Clean-room rule

Customer evidence vs. reference KB

Separates witness evidence from methodology, definitions, false-positive guards, tactics, and confidential reference material.

Source of truthReference onlyNo provenance leakage
Key question
Which material can prove current readiness, and which material may only guide expert interpretation?
Reasoning logic
Use uploaded sources to establish present-state claims. Use criteria, anti-patterns, tactics, taxonomies, and remote PDFs only as controlled analytical context.
Grounding basis
Knowledge taxonomy, allowed and forbidden KB uses, runtime KB index, fallback status, and reference-leak sanitation.
Reliability control
The KB cannot supply a customer quote, justify a score, fill a missing document, or appear as cited provenance in the generated report.
Illustrative thinking sample“The KB defines what embedded AI lifecycle control looks like. It does not prove that the assessed organization has implemented it.”

2. Source registry, chunking & five domain packets

The complete source pack is transformed into traceable chunks and routed into bounded A–E evidence packets before parallel audit begins.

Context is deliberately bounded
Source IDs Page / row / chunk IDs A–E relevance routing 35k target / 45k hard cap
ID
Open logic
Traceability layer

Source registry & provenance

Assigns stable source, page, row, image, and chunk identifiers to every evidence unit.

Source IDChunk IDManifest
Key question
Can every later quote, downgrade, finding, and report claim be traced to an exact source location?
Reasoning logic
Build provenance before interpretation so traceability is structural rather than added after the report is written.
Grounding basis
Document boundaries, PDF page IDs, row numbers, character spans, image metadata, and packet manifests.
Reliability control
A plausible statement without an attributable source location cannot become strong evidence simply because it fits the rubric.
Illustrative thinking sample“The verifier should be able to point to src-004, page 011, chunk 3—not only to a generic ‘AI Strategy’ document.”
A–E
Open logic
Deterministic routing

Domain relevance classification

Scores each source chunk against domain-specific operating, architecture, governance, data, and service signals.

High / medium / lowRouting reasonsPre-classification
Key question
Which chunks are materially relevant to each AI Transformation domain?
Reasoning logic
Use deterministic term-based signals to rank evidence and reduce irrelevant context before the forensic model receives it.
Grounding basis
Domain vocabularies, source names, gap terms, contradiction terms, and relevance tiers.
Reliability control
Each domain receives a tailored evidence universe instead of the full source dump, reducing cross-domain contamination and false pattern matching.
Illustrative thinking sample“Model-release and observability evidence routes primarily to B; service-blueprint evidence routes to E; value hypotheses and portfolio decisions route to C.”
P
Open logic
Bounded evidence packet

Packet construction & weak coverage

Selects the highest-value chunks for each domain and makes thin evidence coverage visible.

Bounded packetGap signalsWeak-coverage flag
Key question
What is the smallest sufficiently rich evidence set each domain auditor should receive?
Reasoning logic
Prioritize high and medium relevance, retain gap and contradiction signals, and stop near the target before context becomes noisy.
Grounding basis
Routing tiers, packet character budgets, included-chunk manifests, coverage notes, and routed images.
Reliability control
Thin packets are labeled as weak coverage. The system does not reinterpret packet silence as healthy maturity or tested anti-pattern absence.
Illustrative thinking sample“A Domain D packet containing only a generic data strategy cannot certify semantic quality, lineage, retrieval quality, or governed AI access.”
RT
Open logic
Execution control

Task-fit model routing & fallback

Maps each task to a controlled primary model profile and ordered fallback chain.

Stage IDsTask fitRun trace
Key question
Which model capability is appropriate for pre-flight, forensic audit, verification, targeted rescan, synthesis, fact-check, and gate explanation?
Reasoning logic
Use different OpenAI and Anthropic profiles according to task shape and consequence, with ordered fallbacks rather than free model choice.
Grounding basis
Central stage registry, reasoning-effort profiles, token limits, provider fallbacks, and recorded actual model use.
Reliability control
Models do not choose their own role. Routing policy controls each task and preserves an auditable execution path.
Illustrative thinking sample“A targeted rescan uses a deeper challenge profile than initial extraction because it must resolve a disputed finding rather than repeat the first pass.”

3. Five parallel forensic domain audits

Each domain audits five maturity criteria and five opposing anti-patterns. Capability evidence, harmful patterns, and missing evidence remain separate throughout the assessment.

Parallel dual-stream reasoning
Shared audit contract: 0–3 evidence scale, exact quotes, one dominant evidence category per quote, no benefit of the doubt, and explicit silence. Across five domains, the Engine evaluates 25 maturity criteria, 25 anti-patterns, and 150 evidence questions.
A
Open logic
Forensic domain

Adaptive Operating Model

Tests decision rhythm, demand routing, Launch-and-Learn, shared learning, and human-AI work redesign.

Decision rightsLearning loopsWork redesign
Expanded sample model
Key questions
Can the organization route different AI work appropriately? Are owners and decisions clear? Does pilot learning become reusable capability? Is work redesigned rather than merely accelerated?
Reasoning logic
Compare operating maturity with decision fog, one-size-fits-all delivery, pilot purgatory, hero culture, and digital Taylorism / workslop.
Grounding basis
Operating cadence, decision-rights maps, demand intake, pilot retrospectives, learning structures, role redesign, and adoption evidence.
Reliability control
An innovation programme or pilot portfolio does not prove an adaptive operating model. Higher scores require recurring decision, learning, ownership, and redesign mechanisms.
Critical note for accuracy
AI activity can increase while transformation capability remains weak. The audit therefore distinguishes more experiments from better organizational absorption.
Illustrative thinking sample“Multiple pilots plus unclear transition to service ownership indicates activity, but may still support Pilot Purgatory rather than embedded readiness.”
B
Open logic
Forensic domain

Enterprise AI Architecture & Platform Readiness

Tests integration, lifecycle control, observability, trust and safety, and AI platform-as-product capability.

ArchitectureLifecycleSafety
Key question
Can AI-enabled services be integrated, released, observed, secured, and reused reliably in the existing enterprise environment?
Reasoning logic
Compare reliable boundaries and lifecycle mechanisms with legacy labyrinths, prompt-to-production chaos, black-box operations, safety theater, and fragmented hidden factories.
Grounding basis
Service interfaces, versioning, release controls, evaluation, model routing, token/model spend, guardrails, red-team results, and platform-product evidence.
Reliability control
A model registry, gateway, or platform purchase is not readiness by itself. The audit requires operating use, traceability, evaluation, ownership, and control evidence.
Illustrative thinking sample“A model gateway proves shared access, but without routing policy, observability, budget controls, and service ownership it does not prove an AI platform product.”
C
Open logic
Forensic domain

AI Strategy, Governance & Value Realization

Tests purpose, impact framing, embedded governance, responsible-AI control, and evidence-based investment.

ValueGovernancePortfolio
Key question
Is AI connected to strategic choices, measurable impact, risk-based control, and evidence-based portfolio decisions?
Reasoning logic
Compare purposeful value creation with slogan strategy, use-case chasing, vanity benefits, rigid gatekeeping, ambiguous accountability, and investment drift.
Grounding basis
Impact Statements, baselines, value hypotheses, governance records, risk classification, budget guardrails, kill/continue/scale decisions, and measured outcomes.
Reliability control
AI spend dashboards, platform usage, ROI language, or efficiency claims support readiness only when tied to ownership, decisions, unit economics, guardrails, and measurable value.
Illustrative thinking sample“A large AI portfolio can indicate strategic activity while simultaneously supporting AI Investment Drift if weak initiatives are not stopped or reprioritized.”
D
Open logic
Forensic domain

Data Foundations, Ownership & Accessibility

Tests domain ownership, semantic context, quality and lineage, governed access, and reusable retrieval patterns.

Data productsSemantic contextRetrieval
Key question
Does AI receive governed, meaningful, traceable, fresh, and reusable data rather than merely available raw data?
Reasoning logic
Compare AI-ready data foundations with ownerless data, context loss, raw data without meaning, broken lineage, leakage, and duplicate knowledge stores.
Grounding basis
Data ownership, semantic definitions, lineage, freshness, access controls, privacy, RAG/retrieval quality, embedding stores, and source coverage.
Reliability control
A data catalog or vector database does not prove semantic quality, ownership, retrieval accuracy, or controlled lifecycle use.
Illustrative thinking sample“A RAG index may exist, but stale sources, duplicate stores, weak ownership, and unbounded context cost still indicate an access-architecture anti-pattern.”
E
Open logic
Forensic domain

Business Capability & Service Architecture

Tests capability anchoring, service blueprints, solution traceability, service-area ownership, and phased scaling.

Service architectureValue streamsScaling
Key question
Are AI initiatives connected to customer needs, service flows, business capabilities, dependencies, and accountable service-area ownership?
Reasoning logic
Compare service-anchored transformation with detached AI ideas, silo automation, untraceable build logic, disconnected project teams, and big-bang scaling.
Grounding basis
Service catalogs, capability maps, customer journeys, service blueprints, value streams, traceability, service-area teams, cost-to-serve, and rollout evidence.
Reliability control
A use-case list does not prove business architecture alignment. Higher maturity requires explicit links to services, outcomes, owners, data, platform dependencies, and learning-based scaling.
Illustrative thinking sample“Automating one task may improve local speed while worsening downstream rework; the audit evaluates the end-to-end service flow, not only the AI feature.”

4. Independent evidence check, anti-pattern semantics & targeted rescan

The first audit is provisional. A separate verification layer tests whether forwarded scores and quotations are genuinely supported before any metric is calculated.

The audit checks itself
Provisional finding Supported / weak / unsupported / missing Max 3 targeted rescans / batch Verified count applied
V
Open logic
Independent verifier

Claim and score verification

Tests each forwarded criterion against the raw packet and exact evidence location.

SupportedWeakUnsupported / missing
Key question
Does the source support both the finding and the strength of the assigned 0–3 score?
Reasoning logic
Verify the quote, surrounding context, score calibration, and packet coverage independently from the forensic auditor.
Grounding basis
Raw source chunks, exact IDs, original and verified counts, quote support, coverage reason, and optional full-source fallback.
Reliability control
The verifier can lower a fluent but over-strong interpretation before metrics are calculated.
Illustrative thinking sample“The source says an AI governance forum is planned. The scanner scored embedded governance at 3. Verify as aspirational and lower the count.”
Ø
Open logic
Absence semantics

Anti-pattern adjudication

Separates confirmed presence, partial presence, tested absence, and unknown absence.

Confirmed presentPartially presentTested / unknown absent
Expanded reliability model
Key questions
Is the harmful pattern present? Is the signal partial? Did the material genuinely cover the risk area well enough to prove absence? Or is absence simply unknown?
Reasoning logic
Give negative findings their own evidence semantics so a zero anti-pattern score cannot automatically become positive readiness.
Grounding basis
Verifier status, original and verified counts, coverage reason, evidence support, and deterministic adjudication rules.
Reliability control
Tested absence is forbidden when the source is silent, weak, contradictory, irrelevant, or already contains a partial harmful signal.
Critical note for accuracy
This prevents the common assessment error of treating missing evidence of a problem as evidence that the problem does not exist.
Illustrative thinking sample“No mention of shadow AI is not tested absence. A comprehensive usage-control review showing governed access and monitored exceptions may support tested absence.”
Open logic
Second opinion

Targeted rescan

Re-examines only the highest-value weak, unsupported, or missing criteria instead of rerunning an entire domain blindly.

Disputed criteriaRescan budgetFocused feedback
Key question
Can a narrower and deeper second pass locate evidence or correct an over-strong first interpretation?
Reasoning logic
Prioritize the largest score deltas and most consequential weak items, rescan at most three criteria per batch, then run verification again.
Grounding basis
Original packet, verifier rationale, original count, status severity, stream, criterion ID, and dedicated rescan routing.
Reliability control
A rescan is not a vote that automatically restores the first score. The second result must still pass independent verification.
Illustrative thinking sample“Re-open B2 and B4 because related controls exist but the evidence is ambiguous; do not regenerate all ten Domain B findings.”
Open logic
Deterministic correction

Apply verified counts

Writes evidence-check outcomes, absence states, adjustment reasons, and rescan status back into Phase 1 before scoring.

Original vs. verifiedAdjustment reasonNo optimism carryover
Key question
What exact count and absence status is Phase 2 permitted to use?
Reasoning logic
Replace unsupported scores, preserve the adjustment trail, and merge all batch verification results into the governed audit record.
Grounding basis
Evidence-check items, adjudication result, adjustment records, and merged batch summaries.
Reliability control
Metrics are calculated from verified counts—not from the initial model’s preferred interpretation.
Illustrative thinking sample“Scanner score 3, verifier score 1, targeted rescan still weak: Phase 2 receives 1 and retains the downgrade reason.”

5. Deterministic metric firewall & confidence bracket

AI does not decide the headline readiness result. Arithmetic converts the verified audit into bounded metrics, classification, and permission for later synthesis.

No generative scoring
Core rule: readiness starts from maturity depth, is reduced by confirmed anti-pattern burden, can receive only a small tested-clearance bonus, and is capped when verified evidence density is sparse.
Σ
Open logic
Metric firewall

Evidence-gated AI readiness

Calculates maturity, burden, clearance, coverage, integrity, density, and A–E domain scores from verified data.

Deterministic math0–100Traceable inputs
Key question
What overall readiness does the verified evidence mathematically permit?
Reasoning logic
Normalize 25 maturity and 25 anti-pattern scores; subtract 50% of anti-pattern burden; add a 10% tested-clearance contribution only when anti-pattern coverage is at least 60%; then apply the evidence cap.
Grounding basis
Verified counts, tested absences, quote-backed gaps, evidence categories, delivered criteria, and domain totals.
Reliability control
No model can round up the result because strategy language, tool adoption, or executive ambition sounds advanced.
Illustrative thinking sample“Strong AI policy evidence plus weak operating proof and entrenched pilot anti-patterns cannot become a high readiness score through narrative synthesis.”
E/S/A
Open logic
Classification

Readiness stage

Translates the capped score and burden into an evidence-aware organizational classification.

InsufficientEmerging / StructuredAdaptive
Key question
What readiness label is defensible after evidence density and harmful patterns are considered?
Reasoning logic
Below the evidence floor: Insufficient evidence. Below 33 readiness: Emerging. Between 33 and 65: Structured, or Scaling with friction when burden exceeds 50. At 66 or above: Adaptive / Value-creating.
Grounding basis
AI readiness, anti-pattern burden, and verified evidence density.
Reliability control
Low apparent burden cannot improve the label when anti-pattern coverage is weak; absence may simply be unknown.
Illustrative thinking sample“A readiness score of 54 with anti-pattern burden above 50 is ‘Scaling with friction,’ not a clean Structured state.”
CAP
Open logic
Evidence cap

Evidence density & readiness ceiling

Caps optimistic readiness when too few criteria have verified source coverage.

<30 BLOCK floor<60 warning capCoverage matters
Key question
How much assessment certainty can the source pack carry, independently of the apparent score?
Reasoning logic
Below 30% density, readiness is capped by available evidence and strategy is blocked. Between 30% and 60%, readiness cannot exceed 60 until stronger current-state evidence is supplied.
Grounding basis
Criteria with verified quotes, quote-backed gaps, anti-pattern findings, or meaningfully tested anti-pattern absence.
Reliability control
The Engine refuses to convert sparse documentation into false organizational maturity certainty.
Illustrative thinking sample“A polished AI strategy covering only ten of fifty evidence surfaces cannot justify a high enterprise readiness result.”
H/M/L
Open logic
Synthesis permission

Confidence bracket

Converts density, delivery integrity, and silent areas into HIGH, MEDIUM, or LOW synthesis behavior.

HIGH directiveMEDIUM cautiousLOW findings-only
Permission model
Key questions
Is there enough verified evidence and pipeline completeness to prescribe action? Should the Engine be directive, cautious, or stop at findings and validation needs?
Reasoning logic
HIGH requires at least 70% evidence density, 95% delivery integrity, and no more than 10 silent areas. LOW is triggered below 30% density, below 70% integrity, or above 18 silent areas. Everything between is MEDIUM.
Grounding basis
Deterministic Phase 2 metrics only.
Reliability control
The synthesis model cannot select a more ambitious mode. Evidence permissions are determined before the model call.
Critical note
LOW-confidence runs still produce useful findings, missing-evidence guidance, and a validation plan; they do not receive directive tactics or a fabricated roadmap.
Illustrative thinking sample“The material may suggest real weaknesses, but if evidence is sparse, the correct output is a validation path—not confident implementation advice.”

6. Evidence summary, diagnosis & three executive lenses

The Engine first creates a fact-only assessment summary, then explains the organizational condition, and only then translates the same evidence through different decision-maker lenses.

Explanation after validation
ES
Open logic
Facts-first synthesis

Assessment evidence summary

Creates the non-prescriptive synopsis of classification, metrics, strengths, gaps, anti-patterns, and missing evidence.

Evidence onlyNo tactic IDsNo directives
Key question
What can the report state directly from verified evidence and locked metrics?
Reasoning logic
Build a factual assessment layer before causal interpretation or planning language is introduced.
Grounding basis
Readiness stage, locked metrics, confirmed strengths, gaps, anti-patterns, verified absences, and silent areas.
Reliability control
Case studies, tactics, invented outcomes, and implementation recommendations are excluded from this layer.
Illustrative thinking sample“The report may state that evidence density is 58% and Domain B scores 7/15; it cannot turn those values into an invented productivity or ROI claim.”
D
Open logic
Interpretive layer

AI Transformation diagnosis

Explains the primary bottleneck, root causes, deterministic domain diagnosis, and confidence without yet prescribing the roadmap.

Primary bottleneckRoot causesA–E diagnosis
Expanded sample model
Key questions
What evidenced mechanism best explains the current readiness state? Which gaps are causal rather than merely adjacent? Where does uncertainty remain?
Reasoning logic
Connect verified patterns across operating model, architecture, strategy, data, and service architecture without changing locked metrics or introducing unproven current-state claims.
Grounding basis
Evidence summary, domain scores, anti-pattern findings, verified absences, source gaps, and confidence bracket.
Reliability control
Diagnosis may explain why, but it cannot prescribe how or create new numbers, owners, controls, or maturity claims.
Critical note for accuracy
Cross-domain coherence is useful only when each link remains traceable. Narrative smoothness must not turn separate weak signals into false causal certainty.
Illustrative thinking sample“AI ambition, weak service ownership, unclear data meaning, and pilot-heavy delivery may jointly indicate an absorption bottleneck—but certainty remains limited if operating evidence is thin.”
TL
Open logic
Persona lens

AI Transformation Lead

Reads the evidence through operating-model readiness, change capacity, portfolio learning, and service-area enablement.

Operating modelLearningScale readiness
Key question
Which readiness gaps block safe organizational absorption and repeatable scaling?
Reasoning logic
Translate the same factual diagnosis into transformation language: operating rhythm, demand routing, learning, ownership, and sequencing.
Grounding basis
Locked evidence summary and diagnosis.
Reliability control
The persona changes emphasis and vocabulary, not scores, source truth, or confidence.
Illustrative thinking sample“Emphasize the missing transition from pilot learning to service-area ownership—not a generic change-management programme.”
CTO
Open logic
Persona lens

CIO / CTO / CDAO

Reads the same diagnosis through architecture, platform, data, lifecycle, security, and operational reliability.

PlatformDataRisk control
Key question
Can AI-enabled services be integrated, governed, observed, and operated reliably at enterprise scale?
Reasoning logic
Use technical and architectural language while preserving the same evidence boundaries and avoiding stack-specific assumptions.
Grounding basis
Architecture, data, lifecycle, observability, cost, retrieval, governance, and security evidence.
Reliability control
The technology lens cannot assume a particular platform, cloud, model, vector store, or architecture unless the source establishes it.
Illustrative thinking sample“Describe missing release traceability and evaluation controls; do not assume a specific MLOps stack or vendor.”
BO
Open logic
Persona lens

Business / Service Area Owner

Reads the evidence through customer value, service outcomes, work redesign, adoption, and accountable ownership.

Customer valueService flowAdoption
Key question
Which AI opportunities are tied to real service outcomes, and what evidence is needed before scaling them?
Reasoning logic
Translate the same diagnosis into customer, service, value-stream, ownership, and human-work consequences.
Grounding basis
Impact statements, service blueprints, value streams, role redesign, adoption, quality, and measured outcome evidence.
Reliability control
The business lens cannot invent customer value or adoption merely because the technology appears promising.
Illustrative thinking sample“Ask whether the pilot improves the whole service flow and first-time-right quality—not only whether one task is completed faster.”
Open logic
Adaptive routing

Synthesis escalation

Routes especially complex or high-friction diagnoses to deeper synthesis using deterministic triggers or explicit deep mode.

Complexity triggersDeep modeRecorded reason
Key question
Is the evidence pattern complex enough to justify deeper analytical effort?
Reasoning logic
Escalate when readiness is very low, burden is very high, gaps are numerous, or the user explicitly requests a deeper mode.
Grounding basis
Readiness, burden, gap count, anti-pattern count, maturity class, and user-controlled mode.
Reliability control
Deeper reasoning changes analytical effort, not the source of truth or the confidence permission.
Illustrative thinking sample“A low-readiness estate with cross-domain anti-patterns may justify deeper causal synthesis, but LOW confidence still forbids a directive roadmap.”

7. Planning decision, Ready-and-Adapt roadmap & tactic permissioning

Recommendations are created only after diagnosis and only in the form permitted by the confidence bracket. The roadmap follows Kickstart → Building the System sequencing and remains linked to verified findings.

Strategy is permissioned
H/M/L
Open logic
Roadmap mode

Directive, cautious, or findings-only

Changes the shape of Phase 3 output according to evidence permission.

HIGH roadmapMEDIUM assumptionsLOW validation plan
Key question
What kind of action output is safe at the current confidence level?
Reasoning logic
HIGH can be directive. MEDIUM keeps phase-level confidence and assumptions. LOW returns evidence-backed findings, candidate themes, missing evidence, and a validation plan—with no directive roadmap.
Grounding basis
Precomputed confidence bracket and locked diagnosis.
Reliability control
The system does not force every assessment into a polished transformation roadmap.
Illustrative thinking sample“When evidence is LOW, validate decision rights, service ownership, lifecycle records, and value baselines before prescribing scale.”
GO?
Open logic
Actionability gate

Planning decision

States whether the roadmap is safe to use, conditionally usable, or blocked pending stronger evidence.

GOCONDITIONAL_GONO_GO
Key question
Which actions are safe to execute now, and what evidence is required before the rest becomes actionable?
Reasoning logic
Separate diagnosis from prognosis: a valid diagnosis can coexist with a NO_GO planning decision when the source pack cannot carry implementation advice.
Grounding basis
Evidence summary, diagnosis confidence, roadmap mode, assumptions, and unresolved evidence gaps.
Reliability control
Actionability is explicit. Readers are not left to infer it from the polish of the report.
Illustrative thinking sample“Safe now: validate service ownership and baseline the target flow. Unsafe now: scale agents across service areas before boundaries and evidence are proven.”
4P
Open logic
Transformation sequencing

Four Ready-and-Adapt phases

Preserves Kickstart validation before Building-the-System scaling.

0–3 months3–6 / 6–1212+ months
Expanded roadmap model
Key questions
What must be established first? What should be proven through a bounded service-area vertical slice? Which platform, data, team, governance, and learning mechanisms can then be embedded and scaled?
Reasoning logic
Phase 1 Emerging — Foundation establishes thesis, owners, demand routing, impact statement, and validation evidence. Phase 2 Structured — Integration proves a vertical slice through service blueprint, data forensics, lifecycle, safety, and pilot playbook. Phase 3 Scaling — Embedding establishes service-area teams, platform-as-product, data products, operating rhythm, and shared learning. Phase 4 Adaptive — Continuous builds Sense & Respond, portfolio learning, safe scaling, value realization, and readiness reassessment.
Grounding basis
Locked findings, Ready-and-Adapt strategy contract, tactic activity playbook, prerequisites, artifacts, roles, and acceptance criteria.
Reliability control
All four phase headings remain visible, but actions may be fewer or empty where evidence does not support a grounded recommendation. The roadmap is not padded with generic transformation work.
Critical note for accuracy
The sequencing prevents platform-first or big-bang advice. Scaling follows evidence from a bounded value-and-safety proof, not executive enthusiasm alone.
Illustrative thinking sample“Do not prescribe enterprise AI operating-model redesign before one service-area vertical slice has clarified value, data, lifecycle, safety, ownership, and absorption needs.”
TAC
Open logic
Tactic permissioning

Verified tactic IDs & activity grounding

Requires exact approved tactic IDs only when the locked problem pattern and prerequisites support them.

Exact TAC IDsWhen to useAcceptance criteria
Key question
Does the recommendation match a verified tactic mechanism, and does the assessed evidence actually support its use?
Reasoning logic
Validate IDs structurally, then compare action language and tactic rules against locked findings. Use the activity playbook for roles, artifacts, implementation activities, risks, and acceptance criteria only after the tactic is grounded.
Grounding basis
Tactics database, activity playbook, taxonomy registry, valid-ID pattern, problem-pattern rules, and prerequisite evidence.
Reliability control
The tactic library is not evidence of need. Invalid or mismatched tactic IDs trigger regeneration, replacement, or removal.
Illustrative thinking sample“A learning-community tactic cannot be attached to generic improvement language unless the evidence shows a learning-flow or fragmented-knowledge gap.”

8. Fact-check, sanitation, Quality Gate & audit trail

The complete report is treated as another artifact to verify. Unsupported claims are challenged, corrected, removed, or blocked before the final publication state is assigned.

Trust is earned after generation
FC
Open logic
Independent verification

Summary & roadmap fact-check

Reviews evidence-summary and diagnosis claims separately from planning decisions and roadmap actions.

Separate checksBounded retriesHigh-reasoning escalation
Key question
Which generated claims are supported, misclassified, tactic-hygiene issues, materially unsupported, or unsafe to act on?
Reasoning logic
Run separate claim packets for narrative and roadmap, merge verdicts, track each pass, and escalate when material blockers survive regeneration.
Grounding basis
Source document, audit logs, locked metrics, diagnosis, approved tactics, claim location, and failure-type taxonomy.
Reliability control
A good diagnosis cannot hide an unsupported roadmap, and a strong roadmap does not excuse fabricated organizational facts.
Illustrative thinking sample“The finding that lifecycle control is weak may be supported, while a promised 30% release-efficiency gain remains fabricated and must be removed.”
SAN
Open logic
Active correction

Strategy, privacy & reference sanitation

Removes, rewrites, or quarantines unsupported claims, identifying content, and confidential reference provenance.

RemoveRewriteQuarantine
Key question
Can an unsafe claim be corrected without changing the evidence contract, or must it disappear from reader-facing output?
Reasoning logic
Operate on exact claims and deep strategy locations. Remove unsafe actions, rewrite known metric misuse, scrub names and identifiers, and replace any leaked reference-document provenance with neutral methodology language.
Grounding basis
Fact-check severity, failure type, source location, privacy terms, reference-KB index, and strategy structure.
Reliability control
Sanitation cannot invent replacement evidence. It can only narrow, correct, anonymize, or remove what the evidence does not support.
Illustrative thinking sample“Replace a leaked KB document name with generic methodology language; remove an unsupported named owner; keep the evidence-backed functional gap.”
QG
Open logic
Publication control

GO / WARN / BLOCK Quality Gate

Aggregates evidence density, traceability, verification, fact-checking, silence, coverage, and sanitation into a visible final decision.

GOWARNBLOCK
Final authority
Key questions
Is the score valid? Is the strategy grounded? Are there material unsupported claims? Is the report safe to act on, usable with warnings, or blocked?
Reasoning logic
Block below the evidence floor, when scored items lack traceable evidence, when Phase 3 validation fails, or when material unsupported claims survive. Warn for partial evidence, low anti-pattern coverage, many silent areas, sanitation, or incomplete fact-checking. Otherwise issue GO.
Grounding basis
Phase 1 and Phase 3 validators, evidence check, deterministic metrics, fact-check, sanitation record, and quality thresholds.
Reliability control
An explanatory model may clarify a WARN or BLOCK decision with source quotes, but it cannot change the deterministic decision.
Critical note for use
A BLOCK does not necessarily invalidate every extracted fact or metric. It means the strategy is unsafe to act on until the listed evidence or grounding issues are resolved.
Illustrative thinking sample“Evidence density 28% triggers BLOCK even when the prose is excellent. The system chooses evidence sufficiency over presentation quality.”
TRACE
Open logic
Assurance layer

Run trace, diagnostics & effective confidence

Records the actual pipeline path and ensures later quality failures can downgrade how the report is rendered.

Model traceDiagnosticsEffective bracket
Key question
Can the organization explain how this assessment ran and prevent an initially confident synthesis from remaining directive after a poor fact-check?
Reasoning logic
Persist stage and model usage, source status, evidence adjustments, warnings, sanitation, and import/export diagnostics. If the Quality Gate flips to BLOCK, effective confidence can downgrade to LOW and hide directive language.
Grounding basis
Run ID, stage events, model-result store, source registry status, confidence bracket, Quality Gate, and report diagnostics.
Reliability control
Initial synthesis permission is not permanent. Post-generation verification can reduce the usable authority of the report.
Illustrative thinking sample“A HIGH-bracket roadmap that later receives BLOCK is rendered as effectively LOW-confidence rather than remaining visibly directive.”

Evidence Summary

Facts, metrics, strengths, gaps, anti-patterns, and missing evidence.

AI Transformation Diagnosis

Primary bottleneck, root causes, and deterministic A–E interpretation.

Persona Views

Transformation Lead, CIO / CTO / CDAO, and Service Owner lenses.

Roadmap or Findings Mode

Directive, cautious, or validation-first output according to evidence permission.

Quality Appendix

Evidence checks, sanitation, source gaps, model trace, and GO / WARN / BLOCK state.

Portfolio role: the Scanner can identify a potential transformation need; the AI Transformation Assessment Engine then performs a specialized evidence-gated diagnosis and creates only the value-creation path the source material permits.
Interpretation boundary This visual shows the skeleton of AI Transformation thinking — not the full implementation
Open scope note
The presentation exposes the high-level reasoning architecture: how the Engine separates input safety, source provenance, dual-stream forensic audit, independent evidence checking, deterministic scoring, diagnosis, confidence-bounded planning, tactic permissioning, fact-checking, and publication control. The production system is built around this skeleton through lower-level contracts and operational mechanisms that are intentionally not shown here.

Reasoning implementation

  • Full prompts, role instructions, and criterion-level question wording
  • Exact orchestration, concurrency, retries, timeout handling, and state transitions
  • Structured-output schemas, JSON repair procedures, and step-to-step contracts
  • Model-specific context shaping, reasoning effort, token limits, and fallback policy

Evidence infrastructure

  • Complete source-registry schemas, routing weights, packet manifests, and DLP rules
  • Evidence taxonomy implementation, provenance fields, quote validation, and visual-evidence handling
  • Anti-pattern adjudication logic, support heuristics, and targeted-rescan prioritization
  • Remote Knowledge Base indexing, source boundaries, storage, and internal-result protection

Knowledge & decision systems

  • All 25 maturity criteria, 25 anti-pattern definitions, and 150 evidence questions
  • Complete tactic database, activity playbook, problem-pattern rules, and case material
  • Scoring formulas, thresholds, calibration choices, and validation-rule details
  • Persona instructions, Ready-and-Adapt contracts, strategy guardrails, and taxonomy contents

Operational assurance

  • Authentication, privacy controls, report import/export safeguards, and secret handling
  • Exact fact-check taxonomies, regeneration appendices, sanitation matching, and quality explanations
  • Run-trace schemas, observability, internal model-result storage, and diagnostics
  • Regression tests, benchmark evolution, drift monitoring, and human validation practices
Design principle: begin with the intended decision and model the expert reasoning system around it. AI is used where probabilistic extraction, verification, diagnosis, and synthesis add value; deterministic controls govern source lineage, score calculation, confidence permissions, tactic validity, correction, and publication. The objective is to minimize free model trust—not to maximize autonomous model judgment.