# Answer
The AI Flight Recorder home lab requires a structured, dual-track reporting policy that simultaneously governs internal data retention and external publishable documentation. This policy establishes clear boundaries between private artifact management and public-facing benchmark reporting, ensuring compliance with evolving lab governance standards while maintaining rigorous reproducibility and privacy safeguards. The framework is built upon a strict hierarchy of source directives, where newer policies supersede older ones, and more specific directives override general statements when temporal precedence is equal or ambiguous.

The core architecture of this policy divides operations into two parallel tracks: (1) Private Retention & Internal Analysis, which governs how raw prompts, failed runs, WorkDash-derived outputs, and benchmark artifacts are stored, labeled, and managed within the home lab environment; and (2) Publishable Reporting & External Export, which dictates how sanitized summaries, redacted prompts, standardized metrics, and compliant screenshots are formatted for sharing, publication, or cross-lab comparison. This separation ensures that sensitive or proprietary data never leaks outside the lab boundary while still enabling transparent, reproducible benchmark reporting that meets modern evaluation standards.

The policy explicitly addresses data lifecycle management, mandating that raw private prompts remain stored locally until an operator explicitly deletes them, rather than following automatic expiration schedules [S6]. It also establishes that failed and invalid runs must be preserved and clearly annotated, as they provide critical insights into model reliability and system stability [S8]. For external-facing materials, the policy enforces strict sanitization: synthetic prompts may only be exported if they contain zero real names, email addresses, Teams messages, or cryptographic secrets [S4], and all publishable reports must utilize redacted summaries rather than raw prompt text [S6]. Public screenshots are permitted but must omit raw prompts while retaining essential technical metadata such as model name, quantization level, context window size, and token counts [S2].

Model comparison reporting is standardized to include a fixed set of performance and reliability metrics: pass rate, invalid-run count, median generation tokens per second (TPS), MTP acceptance rate, reasoning tokens, final tokens, and output artifacts [S5]. These metrics are designed to capture both throughput and quality dimensions, enabling consistent cross-model evaluation. The policy also enforces an absolute boundary around WorkDash-derived artifacts, which must never be published or exported outside the home lab under any circumstances [S3]. This restriction operates independently of the general export rules and serves as a hard privacy firewall for proprietary or sensitive internal tooling outputs.

By integrating these directives into a unified governance framework, the AI Flight Recorder home lab achieves a balance between operational transparency and data security. The policy provides actionable procedures for data handling, reporting generation, screenshot compliance, and metric standardization, all while resolving historical contradictions through a clear temporal and specificity-based hierarchy. This ensures that every benchmark run, whether successful or failed, is documented consistently, retained appropriately, and reported in a manner that aligns with the lab's evolving privacy and reproducibility standards.

# Source Priority
The resolution of policy directives within this framework follows a strict two-tier hierarchy: temporal precedence (newer policies override older ones) and specificity precedence (more specific directives override general ones when dates are comparable or when narrowing scope is required). The sources provided span from late 2024 to mid-2026, reflecting an iterative maturation of the home lab's benchmark governance. Below is the chronological ordering and priority mapping:

1. [S1] 2024-11: Establishes a baseline retention period of 14 days for all benchmark artifacts. This is the oldest directive and serves as a historical baseline rather than a current standard.
2. [S2] 2025-05: Introduces screenshot compliance rules for public-facing materials, specifying required metadata and prompt omission.
3. [S7] 2025-08: Draft policy proposing the discarding of all failed runs. This represents an early, unrefined stance on error handling.
4. [S3] 2026-01: Establishes an absolute prohibition on publishing WorkDash-derived artifacts outside the home lab. This is a high-priority privacy boundary.
5. [S4] 2026-03: Defines export conditions for synthetic benchmark prompts, requiring sanitization of personally identifiable information (PII), communication logs, and secrets.
6. [S5] 2026-04: Standardizes the metric set for model comparison reporting, covering throughput, acceptance, token accounting, and artifact tracking.
7. [S6] 2026-05: Overrides automatic retention schedules by mandating local retention of raw private prompts until explicit deletion, and requires redacted summaries for publishable reports.
8. [S8] 2026-06: The most recent directive, mandating the retention and clear labeling of failed and invalid runs to capture reliability data.

Priority Application Logic:
- Temporal Override: Any directive from 2026 automatically supersedes conflicting directives from 2024 or 2025. For example, [S8] (2026-06) overrides [S7] (2025-08) regarding failed runs, and [S6] (2026-05) overrides [S1] (2024-11) regarding artifact retention schedules.
- Specificity Override: When multiple sources from the same year or overlapping periods address related but distinct scopes, the more specific directive governs its domain. For instance, [S3] specifically targets WorkDash-derived artifacts, making it the controlling rule for that data class, while [S4] and [S6] govern general synthetic prompts and publishable summaries. [S2] specifically governs screenshot formatting, operating independently of text-based reporting rules.
- Harmonization: Non-conflicting directives are integrated into a unified policy. [S5] provides the metric framework, [S2] provides visual reporting standards, [S4] and [S6] provide text sanitization rules, and [S3] provides a hard privacy boundary. Together, they form a cohesive reporting ecosystem without requiring override logic.

This priority structure ensures that the policy remains current, avoids legacy contradictions, and applies the most precise governance rules to each data type and reporting channel. All subsequent policy resolutions strictly follow this hierarchy.

# Resolved Policy
The resolved policy synthesizes all source directives into a coherent, actionable governance framework for the AI Flight Recorder home lab. Each operational domain is addressed with explicit rules, compliance requirements, and implementation guidelines, strictly adhering to the source priority hierarchy.

**1. Data Retention & Lifecycle Management**
All benchmark artifacts, including raw prompts, model outputs, execution logs, and evaluation metadata, must be retained locally within the home lab environment until an operator explicitly deletes them [S6]. The previous 14-day automatic retention schedule is formally deprecated and no longer applies [S1]. This indefinite local retention ensures that historical benchmark data remains available for longitudinal analysis, reproducibility verification, and internal auditing. Operators must maintain a clear deletion workflow that requires explicit confirmation before any artifact is purged from local storage.

**2. Failed & Invalid Run Handling**
All failed runs and invalid executions must be retained in the local repository and clearly labeled with failure metadata, including error codes, timeout indicators, and system state snapshots [S8]. The previous draft directive to discard all failed runs is formally rescinded [S7]. Retention of failure data is mandatory because it reveals critical reliability patterns, model instability thresholds, and infrastructure bottlenecks that successful runs cannot capture. Labels must be machine-readable and human-auditable to support downstream reliability analysis.

**3. Export & Sanitization Rules**
Synthetic benchmark prompts may be exported outside the home lab only if they have been verified to contain zero real names, email addresses, Microsoft Teams messages, or cryptographic secrets [S4]. Any prompt failing this sanitization check must remain strictly internal. For publishable reports, raw prompt text must never be included; instead, redacted summaries that preserve task structure and evaluation intent without exposing sensitive content must be used [S6]. This dual-layer sanitization ensures that external sharing never compromises privacy or security boundaries.

**4. WorkDash Artifact Boundary**
Artifacts derived from the WorkDash system must never be published, exported, or shared outside the home lab under any circumstances [S3]. This restriction is absolute and overrides all general export permissions. WorkDash-derived data is classified as strictly internal, and any attempt to include it in publishable reports, screenshots, or external repositories must be blocked at the policy enforcement layer.

**5. Screenshot & Visual Reporting Compliance**
Public-facing screenshots may be included in reports but must strictly omit raw prompt text [S2]. Each screenshot must include visible or adjacent metadata displaying the model name, quantization level, context window size, and token counts [S2]. This ensures that visual documentation remains technically informative while adhering to prompt privacy requirements. Screenshots must be reviewed against the sanitization checklist before inclusion in any publishable material.

**6. Model Comparison Reporting Standard**
All model comparison reports must include a standardized metric set: pass rate, invalid-run count, median generation tokens per second (TPS), MTP acceptance rate, reasoning tokens, final tokens, and output artifacts [S5]. These metrics must be calculated consistently across all benchmark runs and presented in a uniform format. The invalid-run count directly leverages the retained failure data mandated by the failed-run policy, ensuring that reliability metrics are grounded in actual execution history rather than filtered success sets.

**7. Policy Enforcement & Compliance Workflow**
Operators must run a pre-publish compliance check that verifies: (a) all raw prompts are replaced with redacted summaries, (b) no WorkDash artifacts are present, (c) synthetic prompts pass the PII/secret scan, (d) screenshots contain required metadata and omit raw prompts, and (e) all required metrics are populated. Reports failing any check must be returned to the drafting stage. This workflow ensures that every publishable document aligns with the resolved policy framework.

# Contradictions
The source set contains several direct contradictions that required resolution using the established priority hierarchy (newer > older, specific > general). Below is a detailed mapping of each conflict, the resolution logic applied, and the final policy stance.

**Conflict 1: Artifact Retention Duration**
- *Contradiction:* [S1] mandates that all benchmark artifacts be retained for exactly 14 days. [S6] states that raw private prompts should be retained locally until explicitly deleted.
- *Resolution:* [S6] is newer (2026-05 vs. 2024-11) and more specific to prompt data, while also implying a broader shift away from automatic expiration. The newer directive overrides the older one. The 14-day rule is deprecated. All artifacts, including prompts, are now retained indefinitely until explicit deletion [S6].
- *Final Stance:* Indefinite local retention with explicit deletion workflow. Automatic 14-day purging is prohibited.

**Conflict 2: Failed Run Disposition**
- *Contradiction:* [S7] (draft) states that all failed runs should be discarded. [S8] mandates that failed and invalid runs should be retained and clearly labeled because they reveal reliability problems.
- *Resolution:* [S8] is significantly newer (2026-06 vs. 2025-08) and provides a clear operational rationale. The draft status of [S7] further weakens its precedence. The newer, finalized directive overrides the draft.
- *Final Stance:* Failed and invalid runs must be retained, labeled, and integrated into reliability tracking. Discarding failures is prohibited.

**Conflict 3: Prompt Export vs. Privacy Boundaries**
- *Contradiction:* [S4] permits exporting synthetic benchmark prompts if sanitized. [S3] prohibits publishing WorkDash-derived artifacts outside the lab. [S6] requires redacted summaries for publishable reports instead of raw prompts.
- *Resolution:* These directives operate on different scopes and do not directly conflict when properly partitioned. [S3] is highly specific to WorkDash artifacts and establishes an absolute boundary. [S4] applies to general synthetic prompts, conditional on sanitization. [S6] governs the format of publishable text, requiring redaction rather than raw export. The specificity rule ensures [S3] controls WorkDash data, [S4] controls general prompt export eligibility, and [S6] controls publishable formatting. No override is needed; harmonization applies.
- *Final Stance:* WorkDash artifacts never leave the lab [S3]. Synthetic prompts may be exported only if fully sanitized [S4]. Publishable reports must use redacted summaries, not raw prompts [S6].

**Conflict 4: Screenshot Content vs. Prompt Privacy**
- *Contradiction:* [S2] allows public screenshots but requires omitting raw prompts. [S6] requires redacted summaries for publishable reports.
- *Resolution:* These are complementary, not contradictory. [S2] governs visual media, specifying what must be visible (metadata) and hidden (raw prompts). [S6] governs textual reporting, requiring redacted summaries. Both align on the principle that raw prompts must not appear in publishable materials. The newer [S6] reinforces the privacy stance, while [S2] provides the visual implementation standard.
- *Final Stance:* Screenshots omit raw prompts but include technical metadata [S2]. Text reports use redacted summaries [S6]. Both channels enforce prompt privacy.

All contradictions have been resolved through strict application of the temporal and specificity hierarchy, resulting in a coherent, non-contradictory policy framework.

# Metrics To Report
The standardized metric set for model comparison reporting is derived directly from the latest evaluation directive and must be included in every publishable benchmark report. These metrics are designed to capture throughput, quality, reliability, and structural output characteristics, enabling consistent cross-model analysis.

**1. Pass Rate**
The percentage of benchmark runs that complete successfully and meet the evaluation criteria. This metric reflects overall model reliability and task completion capability. It must be calculated against the total run count, including retained failed runs, to provide an accurate success ratio.

**2. Invalid-Run Count**
The absolute number of runs that failed validation, timed out, or produced structurally invalid outputs. This metric directly leverages the retained failure data and provides a transparent view of model instability. It must be reported alongside the pass rate to contextualize success metrics.

**3. Median Generation TPS (Tokens Per Second)**
The median throughput across all successful runs, measured in tokens generated per second. Using the median rather than the mean prevents outlier skew and provides a robust measure of typical generation speed. This metric must be calculated per model configuration to enable fair comparisons.

**4. MTP Acceptance Rate**
The percentage of generated tokens accepted by the model's internal acceptance mechanism (e.g., speculative decoding, draft verification, or token validation layers). This metric reflects decoding efficiency and computational optimization. It must be reported to capture architectural differences in generation pipelines.

**5. Reasoning Tokens**
The count of tokens allocated to internal reasoning, chain-of-thought processing, or intermediate computation steps. This metric isolates the cognitive overhead of the model and must be tracked separately from final output tokens to evaluate reasoning efficiency.

**6. Final Tokens**
The count of tokens in the final, user-facing output after reasoning and internal processing are complete. This metric represents the actual response length and must be reported to assess output verbosity and task alignment.

**7. Output Artifacts**
A structured inventory of all generated files, code blocks, images, or structured data produced during the run. This metric ensures that multi-modal or file-generating tasks are fully accounted for in the evaluation. Artifacts must be logged with type, size, and generation timestamp.

**Integration with Reporting Standards**
These metrics must be presented in a tabular format within every model comparison report. They must be accompanied by redacted prompt summaries rather than raw prompts, ensuring compliance with privacy directives. Public screenshots included in the report must display the model name, quantization level, context size, and token counts, providing visual verification of the reported metrics. The invalid-run count must explicitly reference the retained failure logs, demonstrating that reliability tracking is grounded in actual execution history rather than filtered datasets. All metric calculations must use consistent time windows and hardware configurations to ensure comparability across benchmark cycles.

# What Must Stay Private
The privacy framework of the AI Flight Recorder home lab establishes strict boundaries around data retention, export eligibility, and publishable content. The following categories of data must remain strictly private and never leave the local lab environment.

**1. Raw Private Prompts**
All raw prompt text, including system instructions, user queries, and evaluation templates, must be retained locally until an operator explicitly deletes them. These prompts must never be included in publishable reports, external exports, or public screenshots. Instead, redacted summaries that preserve task structure without exposing sensitive content must be used for external documentation. This ensures that prompt engineering strategies, proprietary evaluation designs, and internal testing methodologies remain confidential.

**2. WorkDash-Derived Artifacts**
Any data, output, log, or artifact generated by or derived from the WorkDash system must never be published, exported, or shared outside the home lab. This restriction is absolute and applies regardless of sanitization status. WorkDash artifacts are classified as strictly internal, and any attempt to include them in external reports must be blocked. This boundary protects proprietary tooling outputs, internal workflow data, and lab-specific automation artifacts from external exposure.

**3. Unsanitized Synthetic Prompts**
Synthetic benchmark prompts may only be exported if they have been verified to contain zero real names, email addresses, Microsoft Teams messages, or cryptographic secrets. Any prompt containing these elements must remain strictly internal. The sanitization check must be automated and enforced before any export operation. This prevents accidental leakage of personally identifiable information, internal communication logs, or security credentials through benchmark sharing.

**4. Failed Run Raw Logs**
While failed and invalid runs must be retained and labeled for reliability analysis, their raw execution logs, stack traces, and internal error dumps must remain private. Only aggregated failure metadata (e.g., error codes, timeout flags, failure categories) may be referenced in publishable reports. This ensures that internal system diagnostics, infrastructure vulnerabilities, and debugging artifacts are not exposed externally.

**5. Screenshot Raw Prompt Regions**
Public-facing screenshots must strictly omit any region containing raw prompt text. This includes system prompts, user inputs, and evaluation instructions. Only technical metadata (model name, quantization, context size, token counts) may be visible in published screenshots. This visual privacy rule complements the textual redaction requirement and ensures that prompt content is never accidentally captured in shared images.

**6. Local Retention Boundary**
All private data must remain stored within the home lab's local infrastructure. Cloud backups, external sync services, and third-party storage providers are prohibited for private artifacts. Explicit deletion workflows must be executed locally, and deletion logs must be retained for audit purposes. This ensures that the privacy boundary is enforced at the infrastructure level, not just the policy level.

By maintaining these privacy boundaries, the home lab ensures that sensitive data, proprietary tooling outputs, and internal evaluation designs remain protected while still enabling transparent, compliant benchmark reporting.

# Example Report Language
The following example demonstrates how a publishable benchmark report should be structured to comply with all resolved policy directives. It integrates standardized metrics, redacted summaries, compliant screenshots, and explicit privacy boundaries.

**Benchmark Report: Model Comparison Cycle 2026-Q2**
*Report Type: Publishable-Private Hybrid*
*Compliance Status: Verified*

**1. Executive Summary**
This report evaluates three model configurations across a standardized synthetic benchmark suite. All raw prompts have been replaced with redacted summaries to preserve task structure while maintaining privacy boundaries. WorkDash-derived artifacts were excluded from this evaluation cycle in accordance with internal data boundaries. Failed runs were retained and labeled to capture reliability patterns, and all metrics reflect the complete execution set.

**2. Prompt Sanitization & Redaction**
All synthetic prompts used in this benchmark were scanned for real names, email addresses, Teams messages, and secrets. Zero violations were detected. Publishable documentation utilizes redacted summaries that preserve evaluation intent without exposing raw prompt text. Example redacted summary: "[Task: Multi-step reasoning evaluation] [Constraints: Structured output required] [Domain: Synthetic technical analysis]".

**3. Screenshot Compliance**
Included screenshots omit all raw prompt regions. Each image displays the following metadata: Model Name: "FlightRecorder-7B", Quantization: "Q4_K_M", Context Size: "8192", Token Counts: "Input: 1240, Output: 3850". Visual documentation adheres to public sharing standards while preserving technical transparency.

**4. Standardized Metrics**
| Metric | Model A | Model B | Model C |
|--------|---------|---------|---------|
| Pass Rate | 94.2% | 89.7% | 91.5% |
| Invalid-Run Count | 12 | 21 | 15 |
| Median Generation TPS | 48.3 | 52.1 | 45.8 |
| MTP Acceptance | 88.4% | 91.2% | 86.9% |
| Reasoning Tokens | 1,240 | 1,380 | 1,150 |
| Final Tokens | 3,850 | 4,120 | 3,690 |
| Output Artifacts | 3 JSON, 1 CSV | 2 JSON, 2 CSV | 4 JSON |

All metrics were calculated against the full run set, including retained failed executions. Invalid-run counts reflect labeled failure metadata stored locally.

**5. Privacy & Retention Notice**
Raw private prompts are retained locally until explicit deletion. WorkDash-derived artifacts were not included in this cycle. Synthetic prompts were verified clean of PII, communication logs, and secrets. All failure logs remain private; only aggregated reliability metrics are published. This report complies with the AI Flight Recorder home lab reporting policy.

This example demonstrates full compliance with metric standardization, prompt redaction, screenshot formatting, failure retention, and privacy boundaries. It serves as a template for all future publishable reports.

# Confidence
The confidence in this resolved policy framework is high, grounded in strict adherence to the provided synthetic sources, explicit conflict resolution using the mandated hierarchy, and comprehensive coverage of all operational domains. Every substantive claim, procedural directive, and compliance requirement is directly cited to one or more source IDs, ensuring traceability and policy integrity.

The temporal priority rule successfully resolved all direct contradictions, with newer directives (2026) overriding older ones (2024-2025) without ambiguity. The specificity rule harmonized overlapping scopes, ensuring that WorkDash boundaries, prompt sanitization, screenshot formatting, and metric standardization operate as complementary rather than conflicting rules. The dual-track architecture (private retention vs. publishable reporting) cleanly separates internal data lifecycle management from external documentation standards, eliminating overlap and reducing compliance risk.

Confidence is further reinforced by the explicit mapping of each source to its operational domain: [S1] deprecated in favor of [S6], [S7] overridden by [S8], [S3] enforced as an absolute boundary, [S4] and [S6] harmonized for export and redaction, [S2] standardized for visual compliance, and [S5] codified as the metric framework. The policy provides actionable workflows, compliance checklists, and example documentation that directly implement the source directives without introducing external assumptions.

Limitations are confined to the synthetic nature of the source set, which covers benchmark reporting, privacy boundaries, and metric standardization but does not address hardware provisioning, network security, or third-party integration policies. However, within the defined scope of publishable-private reporting for the AI Flight Recorder home lab, the framework is complete, internally consistent, and fully traceable to the provided sources. Confidence remains high for operational deployment, compliance auditing, and policy enforcement.