Research Paper

Decision Infrastructure

The systems, processes, and intelligence layers that transform exploration into qualified execution.

Paper

WL-DI-001

Version

0.3

Status

Living Document

Last Updated

July 2026

Discipline

Systems Engineering

Authors

Warren Labs Research

Implementation

Orqis

Citation

WL-2026-DI-03

This is a living research document. Content evolves as our understanding deepens.

Research Status

Theory■■■■■■■■■□
Implementation■■■■■■■□□□
Validation■■■■■■□□□□
Generalization■■□□□□□□□□

Key Contributions

This paper introduces:

  • Decision Infrastructure as a distinct systems discipline
  • Separation of intelligence quality from decision quality
  • Qualification as an architectural layer between generation and execution
  • Canonical seven-stage lifecycle with continuous feedback
  • Six engineering principles for decision infrastructure design
  • Initial implementation and empirical evidence through Orqis

Abstract

This paper argues that execution should be earned through evidence, not assumed through generation. Organizations can now produce candidates for action faster than they can meaningfully evaluate them, creating a structural gap between generation capacity and judgment capacity. We introduce Decision Infrastructure as the systems discipline focused on this gap — the collection of systems (qualification, simulation, allocation intelligence, outcome attribution, and learning) that together determine whether a candidate has earned the right to consume resources. We propose that this layer is architecturally distinct from intelligence systems (which generate predictions) and execution systems (which deploy resources), and that its absence represents a critical vulnerability in any organization where generation is abundant and the cost of unqualified execution is high. We present a canonical architecture, articulate first principles, and describe an initial implementation in the domain of capital allocation.

01

The Problem

Organizations can now generate plausible strategies, analyses, and recommendations faster than they can meaningfully evaluate them. A single researcher can explore more hypotheses in an afternoon than an entire department could evaluate in a quarter. Strategy generation that once required years of domain expertise can produce hundreds of candidates in minutes.

Every improvement in generation increases the burden on judgment. More candidates means more evaluation, more comparison, more evidence required to distinguish the promising from the plausible. Yet most organizations have invested heavily in generation infrastructure — machine learning pipelines, research teams, AI assistants — while leaving their qualification infrastructure essentially unchanged.

Generation and qualification no longer scale together. The cost of producing another candidate is falling toward zero. The cost of determining whether that candidate deserves real resources — through evidence, simulation, and structured evaluation — remains structurally high. This divergence creates what we call the judgment bottleneck.

Observation

The judgment bottleneck is not a human limitation. It is an infrastructure gap. Organizations have sophisticated systems for generating candidates, but primitive systems for determining which candidates have earned the right to execute.

The common response is better prediction: more accurate forecasts, better models, smarter ranking. But prediction alone cannot solve the problem. A model can estimate that a strategy will succeed with 73% confidence. Whether that strategy should actually be deployed depends on evidence sufficiency, environmental conditions, behavioral consistency, risk tolerance, and whether the strategy has proven itself through structured qualification. These are not prediction problems. They are qualification problems. They require different infrastructure.

Engineering Note

Intelligence systems generate predictions. Execution systems deploy resources. Decision Infrastructure determines which predictions have earned execution. It occupies the gap between the other two.

This gap — between generating a possibility and earning the right to act on it — is where Decision Infrastructure operates. The discipline exists because execution should not be assumed. It should be earned.

Generation creates possibilities. Qualification earns the right to act on them.

Observation

Generation capacity is scaling faster than judgment capacity.

Across AI-assisted research, strategy generation, and automated analysis, the cost of producing another candidate is falling faster than the cost of evaluating it. This creates an asymmetric system: generation expands rapidly while qualification remains dependent on manual review, fragmented metrics, or intuition.

Observed internally

Verified Evidence — Production

Generation-to-qualification funnel

The Orqis implementation generates, validates, simulates, and qualifies strategy candidates through a multi-stage pipeline with 14 explicit qualification criteria (10 global + 4 regime-specific). The qualification gate requires: 7+ days paper trading, 5+ trades, 45%+ win rate, positive ROI, max 30% drawdown, no high-severity flags, intelligence score ≥ 40, and avg trade return ≥ 0.3%.

  • 59 strategies currently hold qualification (22 conditionally, 33 globally qualified regime-unproven, 4 other). 201 never qualified. 266 total eligible. 135 have ever achieved qualification — the gap reflects continuous re-evaluation, not permanent status.
  • Qualification is a production gate — strategies cannot reach live deployment without passing all criteria
  • End-to-end funnel: generation → ~31% pass validation → ~7.5% qualify per foundry sweep → 22.2% reach live eligibility → 5.1% deployed to live. Less than 1% of generated strategies reach live deployment.
  • Rejection rate: 77.8% of evaluated candidates rejected. Rejected strategies average −8.7% ROI vs +6.4% for qualified — the gate makes substantively different decisions, not just slower ones.

Production observed

Qualification Evaluator·VerifiedStrategy Intelligence Layer·VerifiedQualification Proof System (live)·Verified

Counterargument

Is this simply a workflow bottleneck?

A workflow bottleneck can often be solved by adding throughput. The judgment bottleneck is different: processing more candidates without improving qualification can increase false confidence and accelerate poor decisions. The missing capability is not faster review alone, but a system that accumulates evidence, applies gates, and preserves learning.

Evidence supports this distinction: the Orqis qualification system rejects 77.8% of candidates (201 of 266), and rejected strategies average −8.7% ROI vs +6.4% for qualified. If qualification were simply a throughput bottleneck, rejected strategies would perform similarly to qualified ones — they do not. The qualification gate is making substantively different decisions, not just slower ones.

Engineering Implication

If generation and qualification scale differently, they should be designed as separate systems with separate objectives, metrics, and failure modes.

Core Thesis

Execution should be earned through evidence, not assumed through generation.

The infrastructure that enforces this principle — through evidence, qualification, and continuous learning — is what we call Decision Infrastructure.

Definition

Decision Infrastructure

noun

A collection of systems responsible for determining whether intelligence deserves execution. Distinct from intelligence systems (which generate predictions) and execution systems (which deploy resources). Composed of qualification, simulation, allocation intelligence, outcome attribution, and learning layers that together transform exploration into qualified execution.

02

Why Decision Infrastructure Exists

The judgment bottleneck described above is not new — but it has never been treated as an infrastructure problem. Most organizations manage judgment through human processes: committees, executive review, analyst evaluation. These work when generation volume is manageable. They fail when generation scales beyond human evaluation capacity, which is precisely what AI-assisted generation enables.

Decision Infrastructure exists because earning execution at scale requires systematic qualification — not faster human review. The same way manufacturing requires process qualification before scaling production, and pharmaceuticals require clinical trials before market deployment, any domain where execution carries meaningful cost requires structured evidence between generation and action.

Organizations that invest only in generation will drown in optionality. Organizations that invest in qualification will compound their advantage, because every qualification decision — whether it admits or rejects — adds structured evidence that improves future decisions.

Implication

Most systems implicitly assume that generation implies readiness for execution. Decision Infrastructure inverts this default: execution is not the consequence of generation. It is earned through evidence.

Figure 1

The Great Inversion

From Generation Economy to Qualification Economy

Without Decision Infrastructure

Candidate Ideas

Executed Decisions

No gate. No evidence.
Everything proceeds.

With Decision Infrastructure

Candidate Ideas

Evidence

Simulation

Qualification

Learn

Evidence

feeds future decisions

Execute

Outcomes

Result

9 Candidate Ideas

2 Executed Decisions

7 Become Evidence

Claim

Execution should be treated as a permission earned through evidence, not as the default consequence of generation.

Conceptual

Verified Evidence — Production

Validation results

This claim has been tested by comparing outcomes of qualified vs never-qualified strategies using trade-level attribution from production paper trading data.

  • ✓ Cohort comparison: +21.9pp win rate spread (p < 0.0001) across 504 qualified trades vs 2,320 never-qualified
  • ✓ Regime robustness: qualification spread positive in all 6 regimes when controlling for intended_regime
  • ✓ Evidence stability: no significant qualification decay within 60 days
  • ✓ Capital protection: $57.3k protected, 68.8% rejection accuracy
  • ✓ Rejection substantiveness: rejected strategies average −8.7% ROI — the gate makes different decisions, not just slower ones
  • ✓ Account-level spread: +14.9% (59 currently qualified at +6.2% vs 201 never-qualified at −8.6%)

Production observed

Qualification Proof System (live)·VerifiedTrade Attribution Batch·VerifiedQualification Evaluator·Verified

Remaining Evidence Gaps

  • Paper-to-live divergence validation (119 live trades — insufficient, expected October 2026 for 500+)
  • Learning loop direct measurement (indirect evidence is directionally supportive but confounded — controlled comparison requires 4+ weekly learning summaries, expected late August 2026)
  • Risk-adjusted return analysis (win rate is robust but mean return is negative — qualification selects for robustness, not alpha. Is this the right tradeoff?)
  • Regime-mismatch loss reduction tracking ($10,398 current qualified mismatch losses — Phase B regime transition response deployed, should decrease over time)
  • Cross-domain generalization (no non-capital-allocation implementation exists yet)

03

Canonical Architecture

If execution must be earned, the question becomes: what does the earning look like? Decision Infrastructure answers this with a canonical architecture — seven layers organized into three regions (discovery, qualification, deployment) that together define what “earning execution” means in practice.

The architecture is deliberately sequential. Candidates cannot bypass evidence gathering. Evidence cannot bypass simulation. Nothing reaches execution without passing through qualification. This constraint is not bureaucratic — it is structural. Just as a compiler enforces type safety regardless of the programmer’s intentions, Decision Infrastructure enforces qualification regardless of the candidate’s apparent promise.

Figure 2

Decision Infrastructure — Canonical Architecture

Seven layers. One continuous feedback loop.

Discovery

01Exploration

ContextCandidate ideas

02Evidence

Candidate ideasScored evidence

03Simulation

EvidenceBehavioral data

04Qualification

Evidence determines whether execution is earned.

Qualified / not qualified decision

Deployment

05Allocation

Qualified setPrioritized allocation

06Execution

Allocation decisionDeployed decision

07Learning

OutcomesImproved models

Continuous feedback to Exploration

Decision Infrastructure architecture: Exploration produces candidates, Evidence scores them, Simulation generates behavioral data. Qualification determines whether evidence warrants execution. Allocation prioritizes qualified decisions, Execution deploys them. Learning transforms outcomes into improved models and feeds continuously back into Exploration.

Figure 2. Each layer transforms structured inputs into outputs for the next layer; Learning closes the loop by feeding outcomes back into Exploration.

Research Note

The seven-layer model is canonical in the sense that each responsibility must exist somewhere in a trustworthy decision system. The implementation boundaries may vary. A smaller system may combine Evidence and Simulation, while a mature system may split Qualification into multiple gates. What should remain invariant is the separation of exploration, judgment, deployment, and learning.

Counterargument

Must every implementation contain exactly seven services?

No. The model defines responsibilities, not deployment topology. A single service may own multiple layers, or one layer may span several services. The architecture fails only when a responsibility is absent, implicit, or allowed to bypass the qualification lifecycle.

04

Architectural Layers

Each layer owns a single responsibility within the architecture. Inputs arrive. A transformation occurs. Outputs leave. No layer attempts to do what another layer does. This separation of concerns is what makes the architecture composable — each layer can be improved independently without destabilizing the others.

The seven layers are not arbitrary. They represent the minimum viable set of transformations required to move from raw possibility to qualified execution while maintaining a continuous learning feedback loop. Removing any single layer creates a structural gap that degrades decision quality in predictable ways.

01Exploration

Generate possibilities from context, evidence, and learned patterns.

Market contextCandidate ideas

If absent: Without exploration, the system has no raw material to qualify.

02Evidence

Gather structured observations through validation and measurement.

Candidate ideasScored evidence

If absent: Without evidence, qualification decisions rely on intuition alone.

03Simulation

Observe behavior in controlled environments before real-world exposure.

Validated candidatesBehavioral signatures

If absent: Without simulation, behavioral risks are discovered only in production.

04Qualification

Determine whether accumulated evidence warrants execution.

Evidence scoresQualified / not qualified decision

If absent: Without qualification, every idea has implicit permission to execute.

05Allocation

Prioritize qualified opportunities under resource constraints.

Qualified candidatesRanked allocation recommendations

If absent: Without allocation, resources are distributed without intelligence.

06Execution

Deploy resources with precision, accountability, and monitoring.

Allocation decisionsActive positions

If absent: Without guarded execution, deployment has no accountability layer.

07Learning

Transform every outcome into improved future decisions.

Execution outcomesUpdated models

If absent: Without learning, the system repeats the same mistakes indefinitely.

Layer 07 feeds Layer 01 — continuous improvement cycle.

Engineering Note

Each layer should have an explicit owner, defined inputs and outputs, failure behavior, and an audit trail. When any layer is implicit or absent, the architecture degrades silently.

05

First Principles

For execution to be reliably earned rather than assumed, the system that enforces qualification must be governed by explicit constraints. The following six principles define the conditions under which Decision Infrastructure operates. They are engineering constraints — not aspirations.

Six engineering constraints

01Execution should be earned.

No idea should move from exploration to action without structured evidence. Generation does not imply readiness.

02Decision quality is independent of outcome.

A high-quality decision can produce a poor outcome. Systems should optimize the process, not chase results.

03Learning should be continuous.

Every outcome — successful or not — generates information that should systematically improve future decisions.

04Intelligence is becoming abundant.

AI has reduced the marginal cost of idea generation toward zero. This is a technological fact, not a hypothesis.

05Judgment is becoming scarce.

The ability to reliably determine which ideas deserve execution is the new bottleneck. This scarcity creates the need for infrastructure.

06Qualification requires infrastructure.

Judgment at scale cannot rely on individual intuition. It requires systematic qualification systems, simulation environments, and evidence frameworks.

Counterargument

Do formal gates suppress innovation or unconventional ideas?

Poorly designed gates can. Decision Infrastructure should not eliminate exploration; it should make exploration inexpensive while making execution accountable. Novelty remains welcome at the exploration layer. The burden of evidence increases only as a candidate approaches irreversible or costly action.

Engineering Implication

Exploration should be permissive. Execution should be selective. Learning should be compulsory.

Exploration should be permissive. Execution should be selective. Learning should be compulsory.

06

Historical Context

The principle that execution should be earned through evidence is not unique to Decision Infrastructure. Across mature engineering disciplines, the same insight has emerged independently: when the cost of unqualified execution is high enough, organizations develop structured qualification infrastructure between generation and deployment. What is new is not the pattern itself — it is the proposal that this pattern deserves to be generalized into reusable infrastructure.

Pharmaceutical Development 1960s–

Phased clinical trials (I → II → III → IV) require progressive evidence accumulation before market deployment. No pharmaceutical reaches patients without passing through each phase, regardless of how promising early results appear. The cost of unqualified deployment — harm to patients — made qualification infrastructure non-negotiable.

Manufacturing 1980s–

Process qualification (IQ/OQ/PQ) proves reliability before scaling production. Every manufacturing process must demonstrate that it produces consistent results under controlled conditions before it is approved for volume production. The cost of unqualified scaling — defective products, recalls, safety incidents — drove the creation of formal qualification gates.

Software Engineering 2000s–

CI/CD pipelines run automated tests, staging environments, and canary deployments before production release. The evolution from waterfall to continuous delivery was fundamentally a story about shortening the feedback loop between creation and evidence. Modern software teams would not consider deploying code that has not passed through automated qualification — yet many of the same organizations deploy strategies, investments, and business decisions with no qualification infrastructure at all.

Aviation 1950s–

Aircraft certification requires thousands of hours of simulation and testing before commercial operation. An independent qualification authority evaluates evidence from every phase. The consequences of unqualified deployment are so severe that the industry accepts years of structured qualification as the minimum standard.

The pattern is consistent: as the cost of failure increases, industries independently develop structured qualification infrastructure between ideation and deployment. Decision Infrastructure proposes that this convergence is not coincidental — it reflects a fundamental architectural requirement that becomes visible whenever the stakes of execution are sufficiently high.

Decision Infrastructure is not a novel invention. It is the generalization of a pattern that already appears across every mature engineering discipline.

Cross-Domain Evidence

These systems are analogues, not proof that every domain should use identical gates. Their relevance is structural: each independently developed a separation between possibility and deployment through staged evidence gathering.

Cross-domain analogue

Equivalence is structural, not procedural.

07

Open Problems

Any honest description of an emerging discipline must include what it does not yet understand. The following questions guide our current investigation. They are open. Our answers are evolving. We share them not as rhetorical devices, but because we believe that articulating the boundaries of current knowledge is itself a form of intellectual progress.

Several of these questions may not have definitive answers — they may instead reveal design tradeoffs that require different responses in different domains. Understanding which questions are answerable and which are inherently contextual is part of the research.

Evidence LevelConceptualSimulationProductionCross-domainPeer Review

What are the minimal viable components of decision infrastructure across domains?

Priority: High · Horizon: 1-2 years · Open

Comparative implementations in at least three non-capital-allocation domains required.

How should qualification confidence decay over time without new evidence?

Priority: High · Horizon: Current · Partially answered

No statistically significant decay within 60 days (85 strategies, 4 age buckets). The 15–30 day cohort performs best, suggesting a brief settling period.

Can decision infrastructure principles generalize beyond capital allocation?

Priority: Critical · Horizon: 2-5 years · Open

Working implementations with explicit exploration, qualification, execution, and learning boundaries in other domains required.

What is the optimal ratio of exploration cost to qualification cost?

Priority: Medium · Horizon: 1-2 years · Directional data

Foundry: 80/20 exploitation/exploration budget. Generation cost ~$0.025/candidate. Qualification cost (7+ days paper trading) far exceeds generation cost, confirming judgment costs dominate.

How should evidence be weighted across different environmental regimes?

Priority: High · Horizon: Current · Partially answered

Win rate spread positive in all 6 regimes when controlling for intended_regime. Low volatility strongest (+40.2pp). Regime-mismatched trades managed at execution layer, not qualification.

Can behavioral signatures outperform traditional factor models in predicting decision quality?

Priority: Medium · Horizon: 1-3 years · Open

Out-of-sample comparison with predefined evaluation criteria required. 260 behavioral assessments available as training data.

08

Current Implementation

The preceding sections describe why execution should be earned and what the earning looks like architecturally. This section describes what happens when the architecture is actually built. Our research is empirical: architectural patterns cannot be validated through theory alone. They require implementation against real outcomes.

Orqis is the first implementation of Decision Infrastructure, applied to capital allocation. Capital allocation was chosen as the initial domain for three reasons. First, outcomes are quantifiable — profit and loss provide unambiguous feedback on decision quality. Second, feedback loops are fast — strategies can be validated through simulation and paper trading within days rather than years. Third, the cost of unqualified execution is immediate and measurable, creating strong incentives for qualification infrastructure.

Through Orqis, every layer of the canonical architecture has been implemented and tested in production: exploration generates strategy candidates from market context, evidence gathers validation data through backtesting, simulation observes behavior through paper trading, qualification determines live eligibility through multi-criteria gates, allocation prioritizes deployment using environmental intelligence, execution manages live positions with continuous monitoring, and learning feeds outcomes back into improved future generation.

Orqis is not the discipline. It is evidence that the discipline works. The patterns observed in this implementation — particularly the compounding effects of continuous learning and the effectiveness of structured qualification gates — inform our broader research into whether Decision Infrastructure generalizes across domains where decisions are consequential and uncertainty is irreducible.

Current Internal Evidence

Verified Evidence — Production

Qualification cohort divergence

Trade-attributed cohort comparison using production paper trading data. 63 qualified strategies (504 trades) compared against 120 never-qualified strategies (2,320 trades). Trades attributed to qualification windows based on open timestamp.

  • Win rate spread: +21.9 percentage points (60.5% qualified vs 38.6% never-qualified, z = 8.94, p < 0.0001)
  • Win rate spread is positive across all 6 market regimes — qualification consistently selects better strategies
  • Strongest regime: low volatility (+40.2pp win rate spread, +$102/trade PnL spread)
  • No evidence of qualification decay within 60 days — signal is stable
  • Account-level qualification spread: +15.1% (134 qualified at +6.4% avg return vs 143 never-qualified at −8.7%)
  • Top 5 qualified performers: equal-weight portfolio +43.6%, 84.8% win rate, 112 trades

Production observed

Qualification Proof System (live)·VerifiedTrade Attribution Batch·VerifiedQualification Evaluator·Verified

Verified Evidence — Production

Behavioral stability and consistency scoring

The allocation intelligence system computes per-strategy assessments across 6 components: regime performance (30%), performance quality (20%), lifecycle health (15%), behavioral consistency (15%), stability (10%), and diversity contribution (10%). 260 assessments computed in production. Hard floors: negative equity-curve Sharpe or non-positive paper return caps recommendation at "hold" regardless of composite score.

  • 260 strategy assessments computed with 18 incorporating micro-backtest divergence data
  • Micro-backtest integration: low divergence (<0.3) adds +10 quality bonus, high divergence (>0.7) applies −15 penalty
  • Confidence threshold: score ≥ 65 for "increase" (priority deployment) recommendation
  • 8 behavioral families classified with mismatch detection against intended family

Production observed

Behavioral Intelligence Brief·VerifiedAllocation Intelligence (live)·VerifiedBehavioral stability (production)·Verified

Verified Evidence — Production

Regime-adjusted qualification effectiveness

Qualification effectiveness is positive across all 6 market regimes when controlling for strategy intended regime. An initial analysis suggested qualification failed in trending_down — this was an artifact of mixing regime-mismatched trades (strategies designed for other regimes that happened to trade during trending_down periods). When controlled, qualification provides capital protection universally.

  • Low volatility: +40.2pp win rate spread, +$102/trade PnL spread (strongest)
  • Ranging: +19.7pp win rate spread, +$43/trade PnL spread
  • Trending up: +24.0pp win rate spread, +$61/trade PnL spread
  • Choppy: +26.4pp win rate spread, +$25/trade PnL spread
  • Trending down (controlled): qualified strategies lose −$14/strategy vs −$350 for never-qualified (+$336 spread)
  • Volatile: +25.4pp win rate spread (small sample: 8 strategies, 12 trades)
  • Regime-mismatched trades still show positive spread: qualified lose −$125/strategy vs −$284 never-qualified

Production observed

Qualification Proof System (live)·VerifiedTrade Attribution Batch·VerifiedMarket Regime Brief·Verified

Verified Evidence — Production

Evidence decay analysis

Measurement of whether qualification predictions degrade over time. 85 qualified strategies analyzed across 4 age buckets. No statistically significant decay pattern detected within 60 days of qualification.

  • 0–14 days: mean return −0.10%, median +0.03% (13 strategies)
  • 15–30 days: mean +0.17%, median +0.15% (24 strategies — best cohort)
  • 31–60 days: mean −0.11%, median +0.01% (35 strategies)
  • 60+ days: mean +0.08%, median −0.12% (13 strategies)
  • Returns oscillate near zero across all buckets — qualification selects for robustness, not alpha

Internal evidence — verified

Qualification Proof System (live)·VerifiedTrade Attribution Batch·Verified

Verified Evidence — Production

False-negative analysis

Systematic analysis of 113 rejected strategies with sufficient trade history (10+ closed paper trades) to assess actual performance.

  • 66 strategies (58%) correctly rejected — average ROI −10.3%
  • 38 strategies (34%) marginally profitable — average ROI +3.4%, failed robustness criteria (win rate, intelligence score, or fee viability)
  • 9 strategies (8%) strong false negatives — average ROI +13.8%, rejected despite profitability
  • All 9 false negatives share a pattern: high return, low win rate (31–44%, below the 45% threshold). They profit through infrequent large wins offsetting frequent small losses.
  • Design tradeoff: the system sacrifices ~8% false negatives to protect against 58% correctly identified losers. This is appropriate for live capital deployment where drawdown tolerance is limited.

Production observed

Qualification Evaluator·VerifiedTrade Attribution Batch·Verified

Internal Evidence — Preliminary

Paper-to-live divergence

6,662 closed paper trades and 119 closed live trades across approximately 3 months. Sufficient for directional analysis but not statistical significance. Known divergence sources: fees (~0.2% round-trip not modeled in paper), slippage, execution lag.

  • 119 live trades — preliminary sample, insufficient for robust comparison
  • Expected to reach 500+ trades (statistical threshold) by approximately October 2026
  • Backtest-to-paper divergence averaged −2% to −6% per trade — live-to-paper divergence likely smaller but unmeasured

Observed internally

Paper vs Live Brief·Pending reviewExecution Engine Specification·Pending review

Internal Evidence — Indirect

Learning loop effectiveness

6 learning loop phases are deployed. A 33-system audit verified feedback loop closure. Outcome attribution is operational with 260 rows and 159 unique decision fingerprints. Direct before/after measurement is inconclusive (3,743 pre vs 62 post, regime confound). However, indirect operational evidence is directionally supportive.

  • Foundry sweep hit rate: ~1 qualified/day (June) → ~6 qualified/day (July) — 6× improvement coinciding with learning summary deployment
  • Fee bleeder generation rate: dropped from 21% to 12.6% after fee_bleeder_rates injected into template generator (7 family×timeframe combos hard-gated)
  • Paper trading winners (top 10 qualified, with indicator patterns and proven regimes) now injected into AI generation prompt via weekly learning summary
  • Auto-optimization pipeline operational: Optuna creates variants from proven strategies through standard qualification pipeline
  • M1 divergence predictor (AUC 0.6441) wired into foundry admission scoring
  • M4 indicator ranker actively adjusting template indicator periods based on per-regime ML rankings

Observed internally

Learning Loop Implementation·VerifiedOutcome Attribution Design·VerifiedIntelligence & Learning Loop Audit·Verified

Limitations

Limitations & Open Questions

  • Qualification selects for robustness (win rate), not profitability (return magnitude). Qualified strategies have worse mean return (−0.78%) than never-qualified (−0.41%) despite much higher win rate. The system prevents frequent losses but does not prevent occasional large losses.
  • Paper-to-live divergence is unmeasured. All qualification evidence comes from paper trading. 105 live trades exist — insufficient for statistical comparison. The entire evidence base depends on paper trading outcomes predicting live outcomes.
  • Learning loop impact is inconclusive. 6 phases deployed, but before/after comparison shows no measurable improvement (sample too small: 62 post-learning strategies vs 3,743 pre-learning). Re-measurement planned after more learning cycles accumulate.
  • Regime-mismatched losses are a known cost of the learning design. Paper strategies trade through all regimes to accumulate data. Qualified strategies caught in the wrong regime lose −$125/strategy. Execution-level protections (regime transition response) manage this at the deployment layer, not the qualification layer.
  • The current implementation is concentrated in capital allocation — a single domain. Cross-domain generality is unproven.
  • The +14.9% qualification spread uses account-level ROI (all trades). The B1 bucket (qualified-window trades only) shows a different return profile. Both are valid measurements but answer different questions about what qualification selects for.
  • Intended regime vs entry regime distinction: performance attribution must distinguish between a strategy's intended regime (design target) and its entry regime (market condition when a trade opened). Paper strategies trade through all regimes to accumulate learning data. Early analysis that mixed these dimensions incorrectly attributed regime-mismatch losses to qualification failure. When properly controlled, qualification effectiveness is positive across all 6 regimes.
  • Peer review has not yet occurred.

09

What Decision Infrastructure Is Not

Defining a new discipline requires defining its boundaries. Decision Infrastructure is frequently confused with adjacent concepts — prediction engines, dashboards, AI systems, execution platforms. These confusions are understandable, because Decision Infrastructure interacts with all of them. But it is none of them.

Not another prediction model.

Decision Infrastructure does not generate predictions — it qualifies them.

Not another dashboard.

Dashboards display information. Decision Infrastructure acts on it.

Not another AI system.

AI generates intelligence. Decision Infrastructure determines what that intelligence deserves.

Not another execution engine.

Execution engines deploy resources. Decision Infrastructure determines when deployment is warranted.

Not another optimization algorithm.

Optimization improves parameters. Decision Infrastructure qualifies whether those parameters should be trusted.

Important Distinction

Not a guarantee of good outcomes.

Decision Infrastructure improves the quality, auditability, and adaptability of the decision process. It cannot eliminate uncertainty or ensure that every qualified decision succeeds. The goal is better processes, not perfect results.

The discipline evolves.
The architecture endures.

Version History

v0.1March 2026

Initial definition and core lifecycle

v0.2May 2026

Expanded lifecycle, regime intelligence integration

v0.3July 2026Current

Behavioral intelligence, outcome attribution, learning systems

v0.4

Planned: Confidence systems, decision graphs

References

Research — Warren Labs | Orqis