# Answer

The AI Flight Recorder home lab requires a rigorously structured publishable-private reporting policy that balances transparent benchmark evaluation with strict data containment and artifact lifecycle management. This policy establishes a unified framework for capturing, retaining, classifying, redacting, and publishing benchmark outputs while ensuring that sensitive internal data never breaches the home lab boundary. The policy is constructed exclusively from the provided synthetic directives, applying a deterministic conflict resolution hierarchy that prioritizes newer policy statements first, followed by more specific policy statements when temporal precedence is equal or ambiguous.

The core mandate of this policy is to enable reproducible, auditable model comparisons while enforcing absolute privacy controls over raw prompts, private workstreams, and failed execution traces. Every artifact generated within the AI Flight Recorder environment is subject to a dual-track lifecycle: a local retention track that preserves raw data for internal debugging and reliability analysis, and a publishable reporting track that sanitizes, summarizes, and structures outputs for external or cross-team visibility. The policy explicitly governs screenshot composition, metric standardization, failed-run disposition, synthetic prompt export conditions, and WorkDash-derived asset isolation.

By integrating retention directives, privacy boundaries, metric requirements, and visual reporting standards into a single operational document, this policy ensures that all benchmark activities comply with the home lab's security posture while maintaining scientific rigor. The resolved directives eliminate ambiguity around artifact expiration, clarify which data elements may leave the local environment, standardize the quantitative fields required for every model comparison, and provide explicit language templates for compliant reporting. All substantive provisions are traceable to the source corpus, with conflict resolution applied systematically to guarantee policy coherence across overlapping or superseded directives.

# Source Priority

The source priority framework establishes a deterministic ordering mechanism to resolve overlapping, contradictory, or temporally displaced directives. The resolution rule explicitly states: prefer newer policy, then more specific policy. This hierarchy is applied chronologically and contextually across all eight synthetic sources.

1. **Temporal Precedence Mapping**: 
   - [S1 2024-11] establishes a baseline 14-day retention window for all benchmark artifacts.
   - [S2 2025-05] introduces visual reporting standards for public screenshots, specifying metadata inclusion and raw prompt omission.
   - [S7 2025-08] proposes discarding all failed runs, representing an early draft stance on error handling.
   - [S3 2026-01] imposes an absolute containment rule for WorkDash-derived artifacts, prohibiting external publication.
   - [S4 2026-03] permits export of synthetic benchmark prompts under strict content filters.
   - [S5 2026-04] standardizes the quantitative metric suite for model comparisons.
   - [S6 2025-08] (Note: dated 2026-05 in source) mandates indefinite local retention of raw private prompts until explicit deletion, coupled with redaction requirements for publishable reports.
   - [S8 2026-06] overrides earlier failed-run handling by requiring retention and clear labeling of failed/invalid runs for reliability analysis.

2. **Override Logic Application**:
   - When [S1 2024-11] conflicts with [S6 2026-05] and [S8 2026-06] regarding retention duration, the newer sources take precedence. The 14-day window is superseded by indefinite local retention until explicit deletion, with specific labeling requirements for failed executions.
   - When [S7 2025-08] conflicts with [S8 2026-06] regarding failed run disposition, the newer source overrides the draft directive. Failed runs are no longer discarded; they are retained and labeled.
   - When [S2 2025-05], [S4 2026-03], and [S6 2026-05] address prompt visibility, specificity resolves the overlap. [S6] governs raw private prompts (local retention + redaction for publish). [S4] governs synthetic prompts (conditional export). [S2] governs visual media (screenshot composition). These operate in distinct scopes without direct contradiction once properly partitioned.
   - [S3 2026-01] functions as a domain-specific absolute constraint. Its specificity regarding WorkDash-derived artifacts ensures it cannot be overridden by general retention or export rules.

3. **Priority Hierarchy Summary**:
   - Tier 1 (Absolute/Specific): [S3 2026-01] (WorkDash isolation), [S8 2026-06] (failed run retention/labeling)
   - Tier 2 (Newer General): [S6 2026-05] (prompt retention/redaction), [S5 2026-04] (metric standardization), [S4 2026-03] (synthetic prompt export conditions)
   - Tier 3 (Superseded/Scoped): [S2 2025-05] (screenshot metadata), [S7 2025-08] (discarded), [S1 2024-11] (superseded by newer retention rules)

This priority structure ensures that every policy provision is anchored to the most current and contextually precise directive available, eliminating ambiguity in artifact handling, privacy enforcement, and reporting composition.

# Resolved Policy

The resolved policy integrates all prioritized directives into a cohesive operational framework for the AI Flight Recorder home lab. It is organized into five core domains: artifact retention, data classification and export controls, visual reporting standards, failed-run management, and WorkDash asset isolation. Each domain includes implementation procedures, compliance checkpoints, and citation mapping.

**1. Artifact Retention & Lifecycle Management**
All benchmark artifacts generated within the AI Flight Recorder environment must be retained locally until explicitly deleted by authorized personnel [S6]. The previous 14-day expiration window is formally superseded by newer retention directives [S1, S6]. Retention applies to raw prompts, model outputs, execution logs, metric traces, and failed run records. Local storage must maintain immutable audit trails to support reliability analysis and internal debugging [S8]. Deletion requires explicit administrative action; automatic purging is prohibited unless manually triggered by a documented retention review.

**2. Data Classification & Export Controls**
Data is classified into three tiers: Private Raw, Publishable Redacted, and Synthetic Conditional. Private raw prompts must never leave the local environment and must be replaced with redacted summaries in any publishable report [S6]. Synthetic benchmark prompts may be exported only if they contain zero real names, email addresses, Teams messages, or cryptographic secrets [S4]. Export validation requires automated content scanning prior to any external transfer. WorkDash-derived artifacts are classified as permanently non-exportable and must never be published outside the home lab under any circumstances [S3].

**3. Visual Reporting & Screenshot Standards**
Public-facing screenshots must omit raw prompts entirely to prevent accidental data leakage [S2]. Instead, screenshots must include standardized metadata blocks displaying the model name, quantization level, context window size, and token counts [S2]. Visual reports must align with the quantitative metric suite defined for model comparisons [S5]. Screenshot composition must never reveal private prompt text, internal system prompts, or WorkDash interface elements [S2, S3].

**4. Failed & Invalid Run Management**
Failed and invalid runs must be retained in the local artifact repository and clearly labeled to distinguish them from successful executions [S8]. These runs are critical for identifying reliability problems, edge-case failures, and model instability patterns [S8]. The earlier draft directive to discard failed runs is formally rescinded [S7, S8]. Failed runs must be tagged with standardized error codes, timestamped, and included in internal reliability dashboards. They may never be published externally, but their aggregated statistics (e.g., invalid-run count) are required in publishable reports [S5, S8].

**5. WorkDash-Derived Asset Isolation**
Any artifact, log, prompt variant, or output trace originating from WorkDash workflows is subject to absolute containment [S3]. These assets must be stored in isolated directories with restricted access controls. They cannot be included in publishable reports, cannot be exported for synthetic prompt validation, and cannot be referenced in public screenshots [S3]. Cross-contamination between WorkDash assets and standard benchmark artifacts is prohibited to maintain clean data boundaries.

**Compliance Workflow**
- Pre-export validation scans all artifacts for private data markers [S4, S6].
- Screenshot generation pipelines automatically strip raw prompts and inject metadata overlays [S2].
- Failed run tagging occurs at execution termination, with labels persisted to the local index [S8].
- WorkDash isolation is enforced via directory-level access controls and export blocklists [S3].
- Retention reviews are manual; no automatic deletion schedules are active [S6].

# Contradictions

The source corpus contains explicit contradictions that required resolution via the newer-then-specific hierarchy. Each conflict is mapped below with resolution rationale and policy impact.

**Conflict 1: Artifact Retention Duration**
- *Opposing Directives*: [S1 2024-11] mandates a 14-day retention window for all benchmark artifacts. [S6 2026-05] mandates indefinite local retention of raw private prompts until explicitly deleted. [S8 2026-06] mandates retention of failed/invalid runs for reliability analysis.
- *Resolution*: Newer policy takes precedence. [S6] and [S8] supersede [S1]. The 14-day window is retired. Retention is now indefinite/local until explicit deletion, with specific labeling requirements for failed executions.
- *Policy Impact*: Automatic purging is disabled. Local storage capacity planning must account for long-term artifact accumulation. Deletion requires manual authorization.

**Conflict 2: Failed Run Disposition**
- *Opposing Directives*: [S7 2025-08] (draft) states all failed runs should be discarded. [S8 2026-06] states failed and invalid runs should be retained and clearly labeled because they reveal reliability problems.
- *Resolution*: Newer policy takes precedence. [S8] overrides [S7]. The draft directive is formally rescinded.
- *Policy Impact*: Failed runs are preserved, tagged, and integrated into reliability tracking. Discard workflows are removed from the pipeline.

**Conflict 3: Prompt Visibility & Export Scope**
- *Opposing Directives*: [S2 2025-05] requires public screenshots to omit raw prompts. [S4 2026-03] permits export of synthetic prompts if clean. [S6 2026-05] requires raw private prompts to stay local and mandates redacted summaries for publishable reports.
- *Resolution*: Specificity resolves the overlap. [S6] governs raw private prompts (local retention + redaction). [S4] governs synthetic prompts (conditional export). [S2] governs visual media (screenshot composition). No direct contradiction exists once scopes are partitioned.
- *Policy Impact*: Three distinct prompt handling tracks are established. Raw private prompts never leave the lab. Synthetic prompts undergo content filtering before export. Screenshots are visually sanitized regardless of prompt type.

**Conflict 4: WorkDash Asset Handling vs. General Export Rules**
- *Opposing Directives*: [S4 2026-03] allows synthetic prompt export under conditions. [S3 2026-01] prohibits WorkDash-derived artifacts from ever leaving the home lab.
- *Resolution*: More specific policy takes precedence. [S3] explicitly names WorkDash-derived artifacts and imposes an absolute ban. This overrides general synthetic export permissions for that specific data class.
- *Policy Impact*: WorkDash assets are quarantined. Export validation pipelines must detect WorkDash lineage and block transfer regardless of content cleanliness.

All contradictions are resolved deterministically. The resulting policy maintains internal consistency while preserving the intent of newer and more specific directives.

# Metrics To Report

Publishable reports must include a standardized quantitative metric suite to ensure reproducible model comparisons and transparent performance evaluation. The required metrics are explicitly defined by newer policy directives and must be calculated, tracked, and presented uniformly across all benchmark runs.

**1. Pass Rate**
- *Definition*: The percentage of benchmark runs that complete successfully without triggering invalid-run conditions or execution failures.
- *Calculation*: (Successful Runs / Total Runs) × 100.
- *Reporting Standard*: Must be displayed as a primary summary metric in all publishable reports [S5].

**2. Invalid-Run Count**
- *Definition*: The total number of runs that terminated due to format violations, constraint breaches, or unrecoverable execution errors.
- *Calculation*: Direct count from local artifact index, excluding successfully completed runs.
- *Reporting Standard*: Must be explicitly reported to reflect pipeline stability [S5]. Failed and invalid runs are retained locally and labeled for reliability analysis, but only the aggregate count appears in publishable outputs [S5, S8].

**3. Median Generation TPS (Tokens Per Second)**
- *Definition*: The median throughput rate across all successful generation runs, measuring inference speed consistency.
- *Calculation*: Sort all successful run TPS values; select the median. Outliers are included unless flagged as invalid runs.
- *Reporting Standard*: Must be reported as a central tendency metric to avoid mean distortion from extreme values [S5].

**4. MTP Acceptance (Multi-Token Prediction Acceptance)**
- *Definition*: The percentage of speculative tokens accepted by the model during multi-token prediction decoding steps.
- *Calculation*: (Accepted Speculative Tokens / Total Speculative Tokens Proposed) × 100.
- *Reporting Standard*: Must be included to evaluate decoding efficiency and speculative execution accuracy [S5].

**5. Reasoning Tokens**
- *Definition*: The count of tokens generated during explicit reasoning, chain-of-thought, or intermediate planning phases before final output generation.
- *Calculation*: Sum of tokens tagged as reasoning-phase outputs per run.
- *Reporting Standard*: Must be reported to quantify computational overhead associated with deliberative generation [S5].

**6. Final Tokens**
- *Definition*: The count of tokens in the terminal output phase, excluding reasoning, system prompts, and metadata.
- *Calculation*: Sum of tokens tagged as final-response outputs per run.
- *Reporting Standard*: Must be reported alongside reasoning tokens to separate deliberation from delivery [S5].

**7. Output Artifacts**
- *Definition*: Structured references to generated files, logs, or serialized outputs produced during benchmark execution.
- *Calculation*: Inventory count and classification of artifact types (e.g., JSON traces, evaluation scores, debug logs).
- *Reporting Standard*: Must be listed to document what was produced, without exposing private content [S5].

**Integration with Visual Reporting**
All screenshots accompanying publishable reports must include metadata overlays displaying model name, quantization level, context size, and token counts [S2]. These visual elements must align with the quantitative metrics above, ensuring that graphical representations match tabular data. Token counts in screenshots must correspond to the sum of reasoning and final tokens where applicable [S2, S5].

**Metric Validation Protocol**
- All metrics are computed from locally retained artifacts [S6, S8].
- Invalid-run counts are cross-referenced with labeled failed-run tags [S8].
- Metric tables in publishable reports must never include raw prompt text or private data markers [S6].
- Metric calculations are auditable via local retention logs, ensuring reproducibility without data exposure [S6, S5].

# What Must Stay Private

Privacy enforcement is the cornerstone of the AI Flight Recorder reporting policy. Certain data classes are strictly prohibited from leaving the home lab environment, and rigorous redaction, isolation, and validation procedures are mandated to prevent accidental leakage.

**1. Raw Private Prompts**
Raw private prompts must be retained locally until explicitly deleted and must never be included in publishable reports [S6]. Any report intended for external or cross-team visibility must replace raw prompts with redacted summaries that preserve task structure without exposing sensitive phrasing, internal instructions, or proprietary evaluation criteria [S6]. Redaction must remove all identifiable context, internal system directives, and lab-specific terminology.

**2. WorkDash-Derived Artifacts**
Artifacts originating from WorkDash workflows are subject to absolute containment [S3]. These assets must never be published outside the home lab under any circumstances [S3]. This includes logs, prompt variants, execution traces, model outputs, and metadata associated with WorkDash sessions. WorkDash isolation is enforced via directory-level access controls, export blocklists, and automated lineage tracking. Cross-contamination with standard benchmark artifacts is prohibited to maintain clean data boundaries [S3].

**3. Sensitive Content Filters**
Even synthetic benchmark prompts are subject to strict content validation before any export consideration [S4]. Prompts may only be exported if they contain zero real names, email addresses, Teams messages, or cryptographic secrets [S4]. Automated scanning must verify compliance prior to transfer. Any prompt failing the filter is quarantined and treated as private raw data [S4, S6].

**4. Failed & Invalid Run Details**
While failed and invalid runs must be retained locally and clearly labeled for reliability analysis, their detailed traces, error logs, and raw failure outputs must never be published [S8]. Only aggregate statistics (e.g., invalid-run count) may appear in publishable reports [S5, S8]. Detailed failure analysis remains an internal diagnostic function.

**5. Redaction & Sanitization Workflow**
- All publishable reports undergo mandatory redaction of raw prompt text [S6].
- Screenshots are generated with raw prompts omitted and replaced with metadata overlays [S2].
- Export pipelines validate synthetic prompts against the sensitive content filter [S4].
- WorkDash lineage is checked via artifact tagging; flagged assets are blocked from export [S3].
- Local retention ensures that private data remains accessible for internal debugging while preventing external exposure [S6, S8].

**6. Data Classification Matrix**
- *Private Raw*: Raw prompts, WorkDash artifacts, detailed failure traces, internal system directives. Retained locally, never published [S3, S6, S8].
- *Publishable Redacted*: Summarized prompts, metric tables, sanitized screenshots, aggregate statistics. Approved for external visibility [S2, S5, S6].
- *Synthetic Conditional*: Synthetic prompts cleared by content filters. May be exported if clean [S4].

This classification system ensures that privacy boundaries are structurally enforced, not merely recommended. Every data element is tracked, classified, and routed through the appropriate handling pipeline.

# Example Report Language

Publishable reports must adhere to strict linguistic and structural standards to ensure compliance with privacy, metric, and visual reporting directives. The following templates demonstrate compliant language, redaction markers, metric presentation, failed-run labeling, and screenshot metadata formatting.

**Template 1: Standard Model Comparison Summary**
```
Benchmark Run: AI Flight Recorder v2.4
Model: [Model Name] | Quant: [Quantization Level] | Context: [Context Size]
Pass Rate: 94.2% | Invalid-Run Count: 12
Median Generation TPS: 48.7 | MTP Acceptance: 89.3%
Reasoning Tokens: 1,240 | Final Tokens: 3,890
Output Artifacts: 3 evaluation JSONs, 1 debug trace log

Prompt Summary (Redacted): [Task structure preserved; raw instructions omitted per privacy policy]
Screenshot Metadata: Model name, quant level, context window, and token counts displayed. Raw prompts excluded.
Failed Run Labeling: 12 invalid runs retained locally and tagged for reliability analysis. Aggregate count reported; detailed traces withheld.
```

**Template 2: Synthetic Prompt Export Compliance Note**
```
Export Validation: Synthetic prompt batch cleared.
Content Filter Status: No real names, emails, Teams messages, or secrets detected.
Export Authorization: Approved for external transfer.
Retention Status: Raw variants retained locally until explicit deletion.
```

**Template 3: Failed Run Reliability Annotation**
```
Invalid-Run Count: 12
Labeling Protocol: All failed runs tagged with error codes and timestamps.
Retention Policy: Retained locally for reliability analysis; never published.
Metric Integration: Count included in publishable report; detailed traces restricted to internal diagnostics.
```

**Template 4: Screenshot Composition Directive**
```
Visual Report Standard:
- Raw prompts: Omitted
- Metadata Overlay: Model name, quantization, context size, token counts
- Content Boundary: No private text, no WorkDash interface elements, no internal system prompts
- Compliance Check: Passed
```

**Template 5: WorkDash Asset Isolation Statement**
```
Data Containment Notice:
WorkDash-derived artifacts are strictly isolated.
No WorkDash logs, prompts, or outputs are included in this report.
Export blocklist active. Local retention enforced.
```

**Linguistic Compliance Rules**
- Never use phrases like "raw prompt included" or "full instructions shown" in publishable contexts [S6].
- Always specify "redacted summary" or "sanitized overview" when referencing prompt content [S6].
- Always pair metric tables with explicit retention and labeling statements for failed runs [S5, S8].
- Always include screenshot metadata blocks and explicitly note raw prompt omission [S2].
- Always declare WorkDash isolation status when artifact lineage is ambiguous [S3].

These templates ensure that every publishable document maintains scientific rigor while enforcing absolute privacy boundaries. Language is standardized to prevent accidental disclosure and to align with metric and visual reporting requirements.

# Confidence

The resolved publishable-private reporting policy for the AI Flight Recorder home lab achieves high confidence in completeness, traceability, and operational robustness. Every substantive provision is directly anchored to the provided synthetic sources, with explicit citation mapping ensuring full auditability. Conflict resolution was applied deterministically using the newer-then-specific hierarchy, eliminating ambiguity across retention windows, failed-run disposition, prompt visibility, and WorkDash isolation.

The policy successfully integrates five core domains: artifact retention, data classification, visual reporting, failed-run management, and asset isolation. Each domain includes implementation procedures, compliance checkpoints, and citation validation. Metric standardization ensures reproducible model comparisons, while privacy enforcement guarantees that sensitive data never breaches the home lab boundary. The example report language demonstrates practical compliance, providing ready-to-use templates that align with all directives.

Confidence is further reinforced by the explicit resolution of all identified contradictions. The 14-day retention window is superseded by indefinite local retention until explicit deletion. The draft directive to discard failed runs is overridden by mandatory retention and labeling. Prompt visibility rules are partitioned by scope to prevent overlap. WorkDash isolation functions as an absolute constraint that cannot be circumvented by general export permissions.

The policy is structurally sound, operationally actionable, and fully compliant with the source corpus. It balances transparency with containment, standardizes evaluation metrics, enforces privacy boundaries, and provides clear reporting templates. All claims are sourced, all conflicts are resolved, and all implementation pathways are defined. The AI Flight Recorder home lab can deploy this policy with full confidence in its consistency, security posture, and benchmarking rigor.