# Answer

The AI Flight Recorder home lab requires a robust, transparent, and secure reporting policy that balances the need for open benchmarking with strict data privacy and security boundaries. The "publishable-private" nature of this policy means that reports can be shared with trusted peers, collaborators, or within the broader home-lab community, but they must never expose sensitive internal data, personal identifiable information (PII), or proprietary secrets. 

To achieve this, the policy is built upon the latest directives from the home lab's governance framework. The core tenets of the policy are as follows:

1. **Transparency in Reliability:** Failed and invalid benchmark runs must not be hidden. They must be retained, clearly labeled, and reported, as they provide critical insights into model reliability and edge-case failures `[S8]`.
2. **Strict Privacy for Raw Prompts:** Raw private prompts must never be published in full. They must be retained locally until explicitly deleted, and any publishable report must use redacted summaries `[S6]`.
3. **Controlled Export of Synthetic Prompts:** Synthetic benchmark prompts may be exported for sharing, but only if they are completely scrubbed of real names, emails, Teams messages, or secrets `[S4]`.
4. **Comprehensive Metric Reporting:** Model comparisons must report a standardized set of metrics, including pass rate, invalid-run count, median generation TPS, MTP acceptance, reasoning tokens, final tokens, and output artifacts `[S5]`.
5. **Zero Tolerance for Internal Tooling Leaks:** Private WorkDash-derived artifacts must never be published outside the home lab `[S3]`.
6. **Redaction Over Omission:** While older guidelines suggested omitting raw prompts from public screenshots, the newer policy dictates that publishable reports should use redacted summaries, providing more context while maintaining security `[S6]`.

This policy ensures that the AI Flight Recorder home lab maintains its reputation for rigorous, honest benchmarking while protecting the privacy and security of its users and infrastructure.

# Source Priority

When conflicts arise between the provided sources, the resolution follows a strict hierarchy: **Newer Policy > More Specific Policy**. This ensures that the most current and contextually relevant rules govern the reporting process.

1. **Date-Based Priority (Newer Wins):**
   * **S8 (2026-06) vs. S7 (2025-08):** S8 explicitly supersedes S7 regarding failed runs. S7 (2025-08) was a draft policy stating that all failed runs should be discarded. However, S8 (2026-06) establishes that failed and invalid runs should be retained and clearly labeled because they reveal reliability problems. Since S8 is newer, it takes precedence.
   * **S6 (2026-05) vs. S2 (2025-05):** S2 (2025-05) stated that public screenshots may omit raw prompts. S6 (2026-05) provides a more nuanced and updated approach, stating that raw private prompts should be retained locally until explicitly deleted, and publishable reports should use redacted summaries. S6 is newer and more specific to the act of publishing reports, so it overrides S2.
   * **S6 (2026-05) vs. S1 (2024-11):** S1 (2024-11) mandated that all benchmark artifacts should be retained for 14 days. S6 (2026-05) specifies that raw private prompts should be retained locally until explicitly deleted. S6 is newer and more specific to raw private prompts, overriding the general 14-day rule for this specific data type.

2. **Specificity-Based Priority (More Specific Wins):**
   * **S4 (2026-03) vs. S2 (2025-05):** S2 (2025-05) broadly stated that public screenshots may omit raw prompts. S4 (2026-03) specifically addresses synthetic benchmark prompts, stating they may be exported if they contain no real names, emails, Teams messages, or secrets. Because S4 is more specific to the type of prompt (synthetic) and the action (export), it takes precedence over the general screenshot rule in S2.
   * **S3 (2026-01) vs. General Artifact Rules:** S3 (2026-01) specifically targets "Private WorkDash-derived artifacts," mandating they never be published outside the home lab. This specific prohibition overrides any general artifact retention or sharing rules.

By applying this priority framework, the resolved policy ensures that the most current and contextually appropriate rules are always in effect.

# Resolved Policy

The resolved policy for the AI Flight Recorder home lab is divided into five key operational areas: Data Retention, Prompt Handling, Reporting Metrics, Failure Handling, and Export/Publishing.

## 1. Data Retention
* **Raw Private Prompts:** Raw private prompts must be retained locally until explicitly deleted by the user or administrator `[S6]`. This overrides the general 14-day artifact retention rule `[S1]`, as raw prompts contain sensitive data that must be preserved for auditability and debugging purposes.
* **General Benchmark Artifacts:** For artifacts that are not raw private prompts, the general rule of retaining benchmark artifacts for 14 days still applies `[S1]`. However, if an artifact is a private WorkDash-derived artifact, it must never be published outside the home lab `[S3]`.

## 2. Prompt Handling
* **Raw Private Prompts:** Raw private prompts must never be published in full. They must be retained locally until explicitly deleted `[S6]`.
* **Synthetic Prompts:** Synthetic benchmark prompts may be exported for sharing, but only if they are completely scrubbed of real names, emails, Teams messages, or secrets `[S4]`. This ensures that synthetic data does not inadvertently leak PII or sensitive information.
* **Publishable Reports:** When creating publishable reports, raw private prompts must be replaced with redacted summaries `[S6]`. This provides context for the benchmark while maintaining privacy.
* **Public Screenshots:** Public screenshots may omit raw prompts `[S2]`, but the newer policy of using redacted summaries in publishable reports takes precedence `[S6]`.

## 3. Reporting Metrics
* **Model Comparisons:** For model comparisons, the following metrics must be reported: pass rate, invalid-run count, median generation TPS, MTP acceptance, reasoning tokens, final tokens, and output artifacts `[S5]`.
* **Output Artifacts:** Output artifacts must be included in the report, but they must be redacted if they contain sensitive information `[S5]`.

## 4. Failure Handling
* **Failed and Invalid Runs:** Failed and invalid runs must be retained and clearly labeled because they reveal reliability problems `[S8]`. This overrides the older draft policy that suggested discarding all failed runs `[S7]`.
* **Invalid-Run Count:** The invalid-run count must be reported as part of the model comparison metrics `[S5]`.

## 5. Export/Publishing
* **Private WorkDash-Derived Artifacts:** Private WorkDash-derived artifacts must never be published outside the home lab `[S3]`.
* **Synthetic Prompts:** Synthetic benchmark prompts may be exported if they contain no real names, emails, Teams messages, or secrets `[S4]`.
* **Publishable Reports:** Publishable reports should use redacted summaries of raw private prompts `[S6]`.

# Contradictions

The following contradictions were identified between the sources and resolved according to the priority rules:

## Contradiction 1: Handling of Failed and Invalid Runs
* **Source S7 (2025-08):** A draft policy stating that all failed runs should be discarded.
* **Source S8 (2026-06):** A policy stating that failed and invalid runs should be retained and clearly labeled because they reveal reliability problems.
* **Resolution:** S8 is newer (2026-06 > 2025-08) and more specific to the purpose of benchmarking (revealing reliability problems). Therefore, S8 supersedes S7. Failed and invalid runs must be retained and labeled.

## Contradiction 2: Handling of Raw Prompts in Public Screenshots
* **Source S2 (2025-05):** Stated that public screenshots may omit raw prompts but should include model name, quant, context size, and token counts.
* **Source S6 (2026-05):** Stated that raw private prompts should be retained locally until explicitly deleted, and publishable reports should use redacted summaries.
* **Resolution:** S6 is newer (2026-05 > 2025-05) and more specific to the act of publishing reports. Therefore, S6 supersedes S2. Publishable reports must use redacted summaries of raw private prompts, rather than simply omitting them.

## Contradiction 3: Retention of Raw Private Prompts
* **Source S1 (2024-11):** Stated that all benchmark artifacts should be retained for 14 days.
* **Source S6 (2026-05):** Stated that raw private prompts should be retained locally until explicitly deleted.
* **Resolution:** S6 is newer (2026-05 > 2024-11) and more specific to raw private prompts. Therefore, S6 supersedes S1. Raw private prompts must be retained until explicitly deleted, overriding the 14-day rule.

## Contradiction 4: Export of Synthetic Prompts vs. Omission in Screenshots
* **Source S2 (2025-05):** Stated that public screenshots may omit raw prompts.
* **Source S4 (2026-03):** Stated that synthetic benchmark prompts may be exported if they contain no real names, emails, Teams messages, or secrets.
* **Resolution:** S4 is newer (2026-03 > 2025-05) and more specific to synthetic prompts and the action of exporting. Therefore, S4 supersedes S2. Synthetic prompts may be exported if they are scrubbed of PII and secrets, rather than simply being omitted from screenshots.

# Metrics To Report

For model comparisons, the following metrics must be reported to ensure comprehensive and transparent benchmarking `[S5]`. Each metric provides critical insights into the model's performance, reliability, and efficiency.

1. **Pass Rate:** The percentage of benchmark runs that successfully completed without errors. This metric indicates the model's overall reliability and ability to follow instructions.
2. **Invalid-Run Count:** The number of runs that failed due to invalid inputs, timeouts, or other errors. This metric highlights edge cases and potential failure modes.
3. **Median Generation TPS (Tokens Per Second):** The median speed at which the model generates tokens. This metric provides a robust measure of performance, as it is less affected by outliers than the mean.
4. **MTP Acceptance (Multi-Token Prediction Acceptance):** The rate at which multi-token predictions are accepted by the model. This metric indicates the efficiency of speculative decoding and the model's ability to predict future tokens.
5. **Reasoning Tokens:** The number of tokens generated during the reasoning phase. This metric provides insights into the model's ability to think through complex problems step-by-step.
6. **Final Tokens:** The total number of tokens generated in the final output. This metric indicates the length and detail of the model's responses.
7. **Output Artifacts:** The actual outputs generated by the model, including text, code, or other structured data. These artifacts must be included in the report but must be redacted if they contain sensitive information.

These metrics must be reported for all model comparisons to ensure that the benchmarking process is transparent, reproducible, and informative.

# What Must Stay Private

To maintain the "publishable-private" nature of the AI Flight Recorder home lab, the following data must never be published outside the home lab:

1. **Raw Private Prompts:** Raw private prompts must never be published in full. They must be retained locally until explicitly deleted, and any publishable report must use redacted summaries `[S6]`.
2. **Private WorkDash-Derived Artifacts:** Private WorkDash-derived artifacts must never be published outside the home lab `[S3]`. This includes any internal tooling, dashboards, or configurations derived from WorkDash.
3. **PII and Secrets in Synthetic Prompts:** Synthetic benchmark prompts may be exported, but only if they are completely scrubbed of real names, emails, Teams messages, or secrets `[S4]`. Any synthetic prompt that contains PII or secrets must not be exported.
4. **Unredacted Output Artifacts:** Output artifacts must be included in the report, but they must be redacted if they contain sensitive information `[S5]`. This includes any PII, secrets, or proprietary information that may have been inadvertently generated by the model.
5. **Internal Infrastructure Details:** Any details about the home lab's infrastructure, such as IP addresses, network configurations, or hardware specifications, must not be published. This is implied by the strict privacy requirements for raw private prompts and WorkDash artifacts.

By strictly adhering to these privacy boundaries, the AI Flight Recorder home lab can share its benchmarking results with the broader community while protecting the privacy and security of its users and infrastructure.

# Example Report Language

The following is an example of how a publishable-private report for the AI Flight Recorder home lab should be structured and written. This example demonstrates how to redact sensitive information, report metrics, and label failed runs.

## AI Flight Recorder Home Lab - Benchmark Report

**Report Date:** 2026-06-15
**Benchmark Suite:** Synthetic Reasoning & Code Generation
**Models Compared:** Model A (Quantized), Model B (Full Precision)

### 1. Executive Summary

This report compares the performance of Model A and Model B on the Synthetic Reasoning & Code Generation benchmark suite. Model A demonstrates superior reasoning capabilities, while Model B shows better code generation accuracy. Both models exhibit high reliability, with minimal invalid runs.

### 2. Metrics Comparison

| Metric | Model A (Quantized) | Model B (Full Precision) |
| :--- | :--- | :--- |
| Pass Rate | 94.2% | 91.8% |
| Invalid-Run Count | 3 | 7 |
| Median Generation TPS | 45.2 | 38.7 |
| MTP Acceptance | 0.82 | 0.79 |
| Reasoning Tokens | 1,245 | 1,180 |
| Final Tokens | 2,890 | 2,750 |
| Output Artifacts | [Redacted] | [Redacted] |

### 3. Prompt Handling and Redaction

**Raw Private Prompts:** Raw private prompts have been redacted for privacy. The following is a redacted summary of the prompts used in this benchmark:

* **Prompt 1:** "Generate a Python function to [REDACTED] given [REDACTED] as input."
* **Prompt 2:** "Explain the concept of [REDACTED] in the context of [REDACTED]."
* **Prompt 3:** "Write a SQL query to [REDACTED] from the [REDACTED] table."

**Synthetic Prompts:** Synthetic prompts used in this benchmark have been scrubbed of PII and secrets. The following is an example of a synthetic prompt:

* **Synthetic Prompt:** "Generate a Python function to sort a list of [REDACTED] by [REDACTED]."

### 4. Failure Analysis

**Failed and Invalid Runs:** Failed and invalid runs have been retained and clearly labeled. The following is a summary of the failures:

* **Model A:**
  * **Run 1:** Invalid run due to timeout. [REDACTED]
  * **Run 2:** Invalid run due to invalid input. [REDACTED]
  * **Run 3:** Failed run due to model hallucination. [REDACTED]
* **Model B:**
  * **Run 1:** Invalid run due to timeout. [REDACTED]
  * **Run 2:** Invalid run due to invalid input. [REDACTED]
  * **Run 3:** Failed run due to model hallucination. [REDACTED]
  * **Run 4:** Failed run due to model hallucination. [REDACTED]
  * **Run 5:** Failed run due to model hallucination. [REDACTED]
  * **Run 6:** Failed run due to model hallucination. [REDACTED]
  * **Run 7:** Failed run due to model hallucination. [REDACTED]

### 5. Output Artifacts

**Model A Output Artifacts:** [Redacted]
**Model B Output Artifacts:** [Redacted]

### 6. Conclusion

Model A demonstrates superior performance in terms of pass rate, median generation TPS, and MTP acceptance. Model B shows better code generation accuracy but has a higher invalid-run count. Both models exhibit high reliability, with minimal failures. The AI Flight Recorder home lab recommends Model A for general-purpose use, while Model B may be better suited for specific code generation tasks.

### 7. Privacy and Security Statement

This report has been generated in accordance with the AI Flight Recorder home lab's publishable-private reporting policy. Raw private prompts have been redacted, and synthetic prompts have been scrubbed of PII and secrets. Private WorkDash-derived artifacts have not been published. All failed and invalid runs have been retained and clearly labeled.

# Confidence

The confidence in this policy is high. The sources provided are explicit and unambiguous, and the conflict resolution rules are clear. The policy is built upon the latest directives from the home lab's governance framework, ensuring that it is current and relevant. The policy balances the need for transparency and reliability with strict privacy and security boundaries, making it suitable for the "publishable-private" nature of the AI Flight Recorder home lab. The detailed examples and explanations provided in this report ensure that the policy is easy to understand and implement.