# Answer
The AI Flight Recorder home lab operates under a comprehensive reporting and data management policy designed to balance rigorous performance analysis with strict privacy and security protocols. This policy governs the lifecycle of benchmark artifacts, the methodology for model comparison, and the specific requirements for public-facing documentation.

**1. Data Retention and Lifecycle Management**
All raw private prompts must be retained locally within the home lab environment until they are explicitly deleted by an authorized user [S6]. While an earlier policy suggested a 14-day retention window for all benchmark artifacts [S1], the more recent and specific instruction regarding raw private prompts [S6] takes precedence. Furthermore, the lab maintains a specific protocol for failed and invalid runs: these must be retained and clearly labeled rather than discarded, as they provide critical insights into reliability problems [S8]. This overrides previous drafts that suggested discarding failed runs [S7].

**2. Publication and Export Standards**
The lab distinguishes between internal artifacts and public-facing reports. Private WorkDash-derived artifacts are strictly prohibited from being published outside the home lab [S3]. For public-facing reports, raw private prompts must never be included; instead, reports must utilize redacted summaries [S6]. 

When publishing screenshots, the following requirements must be met:
- Raw prompts must be omitted [S2].
- The model name must be included [S2].
- The quantization (quant) must be included [S2].
- The context size must be included [S2].
- The token counts must be included [S2].

Synthetic benchmark prompts may be exported provided they are scrubbed of all sensitive information, specifically:
- Real names [S4].
- Emails [S4].
- Teams messages [S4].
- Secrets [S4].

**3. Model Comparison and Reporting Metrics**
When conducting model comparisons, the following metrics are mandatory for every report:
- Pass rate [S5].
- Invalid-run count [S5].
- Median generation TPS (Tokens Per Second) [S5].
- MTP (Maximum Token Probability) acceptance [S5].
- Reasoning tokens [S5].
- Final tokens [S5].
- Output artifacts [S5].

**4. Security and Privacy Safeguards**
The primary directive is the protection of internal data. WorkDash-derived artifacts are classified as strictly private [S3]. Any data intended for export or publication must undergo a verification process to ensure no real names, emails, Teams messages, or secrets are present [S4]. All public-facing content must prioritize redacted summaries over raw data [S6].

# Source Priority
The policy resolution follows a strict hierarchy:
1. **Recency (Newest Policy):** Policies from 2026 (S3, S4, S5, S6, S8) take precedence over those from 2025 (S2, S7) and 2024 (S1).
2. **Specificity:** Where dates are equal or where a policy addresses a specific subset of data (e.g., "raw private prompts" vs. "benchmark artifacts"), the more specific policy is applied.
3. **Conflict Resolution:** 
    - S8 (2026-06) overrides S7 (2025-08) regarding failed runs.
    - S6 (2026-05) overrides S1 (2024-11) regarding prompt retention.
    - S6 (2026-05) overrides S2 (2025-05) regarding the use of redacted summaries in reports.

# Resolved Policy
**I. Artifact Retention Protocol**
- **Raw Private Prompts:** Must be stored locally and retained indefinitely until an explicit deletion command is issued [S6].
- **Failed/Invalid Runs:** Must be preserved and clearly tagged with a "Failure" or "Invalid" label to facilitate reliability debugging [S8].
- **General Benchmark Artifacts:** While a 14-day window was previously suggested [S1], the current mandate for raw prompt retention [S6] and failure retention [S8] establishes a "retain until deleted" standard for primary data.

**II. Publication and Export Rules**
- **Internal Only:** Any artifact derived from WorkDash is strictly confined to the home lab [S3].
- **Public Screenshots:** Must contain: Model Name, Quantization, Context Size, and Token Counts. They must *not* contain raw prompts [S2].
- **Public Reports:** Must use redacted summaries of prompts rather than the raw input [S6].
- **Synthetic Prompt Export:** Permitted only if the prompt is verified to contain zero real names, emails, Teams messages, or secrets [S4].

**III. Comparative Analysis Requirements**
Every model comparison report must include a standardized metrics block containing:
- Pass Rate [S5].
- Invalid-run Count [S5].
- Median Generation TPS [S5].
- MTP Acceptance [S5].
- Reasoning Tokens [S5].
- Final Tokens [S5].
- Output Artifacts [S5].

# Contradictions
The following contradictions were identified and resolved during the policy construction:
1. **Retention of Failed Runs:** S7 (2025-08) stated that failed runs should be discarded. This was overruled by S8 (2026-06), which mandates their retention and labeling to identify reliability issues. S8 is newer and therefore takes precedence.
2. **Retention Duration:** S1 (2024-11) suggested a 14-day retention for all benchmark artifacts. This was overruled by S6 (2026-05), which specifies that raw private prompts must be retained until explicitly deleted. S6 is newer and more specific.
3. **Reporting Style:** S2 (2025-05) focuses on the content of screenshots (omitting prompts but including metadata). S6 (2026-05) provides a broader reporting standard requiring redacted summaries for publishable reports. S6 is newer and takes precedence for the general reporting structure.

# Metrics To Report
The following metrics are mandatory for all model comparisons [S5]:
1. **Pass Rate:** The percentage of successful completions.
2. **Invalid-run Count:** The total number of runs that failed or produced invalid outputs.
3. **Median Generation TPS:** The median speed of token generation.
4. **MTP Acceptance:** The acceptance rate of Maximum Token Probability.
5. **Reasoning Tokens:** The count of tokens generated during the reasoning phase.
6. **Final Tokens:** The count of tokens in the final output.
7. **Output Artifacts:** The actual files or data structures produced by the model.

# What Must Stay Private
The following items are strictly prohibited from being published outside the home lab:
- **WorkDash-derived Artifacts:** Any data or artifacts originating from WorkDash [S3].
- **Raw Private Prompts:** These must remain local; only redacted summaries are allowed in public reports [S6].
- **Personally Identifiable Information (PII):** Real names and emails [S4].
- **Internal Communications:** Teams messages [S4].
- **Security Credentials:** Any secrets [S4].

# Example Report Language
**Example 1: Model Comparison Summary (Public)**
"Model Comparison: [Model Name] ([Quant])
Context Size: [Context Size]
Metrics:
- Pass Rate: [X]%
- Invalid-run Count: [Y]
- Median Generation TPS: [Z]
- MTP Acceptance: [A]%
- Reasoning Tokens: [B]
- Final Tokens: [C]
- Output Artifacts: [Description of artifacts]

*Note: Prompt details have been redacted for privacy [S6].*

**Example 2: Screenshot Caption (Public)**
"Screenshot of [Model Name] ([Quant]) at [Context Size] context. Total Token Count: [Count]. (Raw prompt omitted per policy [S2])."

**Example 3: Internal Failure Log (Private)**
"Run ID: [ID] - Status: INVALID/FAILED
Reason: [Description]
Note: Retained for reliability analysis [S8].
Source: WorkDash-derived artifact [S3] - DO NOT PUBLISH."

# Confidence
High. The policy was constructed by strictly adhering to the provided sources, resolving all chronological and specificity conflicts as instructed. All metrics and privacy constraints are directly mapped to the provided synthetic sources.

**Detailed Procedural Guidelines for Lab Technicians**

To ensure consistent adherence to the AI Flight Recorder home lab policies, all personnel must follow these standardized procedures when handling data, generating reports, and managing benchmark artifacts.

**I. Data Ingestion and Retention Procedures**
1. **Initial Capture:** Upon the generation of any raw private prompt, the system must automatically flag the data for local retention. According to policy [S6], these prompts are to be stored in the local lab environment indefinitely. They must not be moved to cloud storage or shared with external entities.
2. **Retention Logic:** While historical policies suggested a 14-day lifecycle for all artifacts [S1], the current mandate [S6] requires that raw private prompts remain accessible until an authorized user performs an explicit deletion. Technicians should ensure that the "Delete" command is only issued after a prompt has been fully analyzed and its data has been successfully transitioned into a redacted summary for public reporting.
3. **Failure Handling:** When a model run results in an "Invalid" or "Failed" status, the technician must not discard the output. Per policy [S8], these runs are critical for identifying reliability problems. The technician must:
    - Retain the full output of the failed run.
    - Apply a clear, high-visibility label (e.g., "FAILURE_LOG" or "INVALID_RUN").
    - Document the specific reason for the failure in the internal metadata.
    - This procedure directly supersedes the outdated draft [S7] which suggested discarding failed runs.

**II. Public Reporting and Screenshot Preparation**
1. **Redaction Protocol:** Before any report is moved to a "Publishable" status, the technician must review the content to ensure no raw private prompts are present. Instead, the technician must provide a "Redacted Summary." This summary should describe the intent and general parameters of the prompt without revealing the specific, sensitive input [S6].
2. **Screenshot Metadata Requirements:** When capturing screenshots for public consumption, the following metadata must be overlaid or included in the caption:
    - **Model Name:** (e.g., "Llama-3-70B") [S2].
    - **Quantization:** (e.g., "4-bit", "FP16") [S2].
    - **Context Size:** (e.g., "32k", "128k") [S2].
    - **Token Counts:** Total input and output tokens [S2].
    - **Prompt Exclusion:** The raw prompt must be completely omitted from the visual frame [S2].
3. **WorkDash Integrity:** Technicians must verify that no artifacts derived from WorkDash are included in any public-facing documentation. WorkDash-derived artifacts are strictly classified as "Internal Only" [S3].

**III. Synthetic Prompt Export Workflow**
If a technician needs to export synthetic benchmark prompts for external use, they must perform a four-point security scrub:
1. **Name Scrub:** Verify that no real names are present in the prompt text [S4].
2. **Email Scrub:** Verify that no email addresses are present [S4].
3. **Communication Scrub:** Verify that no Teams messages or internal chat logs are embedded [S4].
4. **Secret Scrub:** Verify that no API keys, passwords, or other secrets are present [S4].
Only after all four checks are passed can the synthetic prompt be exported.

**Data Classification and Handling Matrix**

To simplify compliance, the following matrix defines how different data types should be handled within the AI Flight Recorder home lab:

| Data Type | Source Policy | Retention Rule | Publication Status | Handling Instructions |
| :--- | :--- | :--- | :--- | :--- |
| **WorkDash Artifacts** | [S3] | Local Only | **Strictly Prohibited** | Never publish outside the home lab. |
| **Raw Private Prompts** | [S6] | Retain until deleted | **Redacted Only** | Keep local; use redacted summaries for reports. |
| **Failed/Invalid Runs** | [S8] | Retain & Label | **Internal Only** | Label clearly to reveal reliability problems. |
| **Synthetic Prompts** | [S4] | Exportable | **Conditional** | Export only if scrubbed of names, emails, Teams msgs, secrets. |
| **Public Screenshots** | [S2] | N/A | **Public** | Must include Model, Quant, Context, and Tokens; omit raw prompts. |
| **Model Comparison Data**| [S5] | N/A | **Public** | Must include all 7 mandatory metrics. |

**Metric Definitions and Significance**

For every model comparison, the following metrics [S5] must be reported. Each serves a specific purpose in the AI Flight Recorder's evaluation framework:

1. **Pass Rate:** Measures the overall success of the model in completing the assigned tasks. This is the primary indicator of utility.
2. **Invalid-run Count:** Quantifies the frequency of system errors, timeouts, or non-responsive outputs. This is essential for identifying infrastructure or model stability issues.
3. **Median Generation TPS:** Provides a standardized measure of the model's inference speed. Using the median (rather than the mean) ensures that outliers do not skew the perceived performance.
4. **MTP Acceptance:** Measures the Maximum Token Probability acceptance rate, which helps determine the model's confidence and adherence to probability thresholds during generation.
5. **Reasoning Tokens:** Tracks the number of tokens generated during the "Chain of Thought" or internal reasoning phase. This helps differentiate between "thinking" time and "output" time.
6. **Final Tokens:** Tracks the length of the final response provided to the user, useful for analyzing output conciseness and efficiency.
7. **Output Artifacts:** Documents the actual files, code blocks, or data structures produced, providing a qualitative look at the model's capabilities.

**Privacy and Security Audit Protocol**

Before any data leaves the home lab or is moved to a public-facing repository, it must pass the "Flight Recorder Privacy Audit." The auditor must check for the following:

- **WorkDash Leakage Check:** Does the artifact contain any WorkDash-derived data? If yes, it must be quarantined [S3].
- **PII Scrubbing:** Does the report contain real names or email addresses? If yes, it must be redacted [S4].
- **Internal Comms Check:** Does the report contain Teams messages? If yes, it must be removed [S4].
- **Secret Check:** Does the report contain any secrets or credentials? If yes, it must be removed [S4].
- **Prompt Redaction Check:** Does the public report contain a raw private prompt? If yes, it must be replaced with a redacted summary [S6].
- **Screenshot Metadata Check:** Does the screenshot include the Model Name, Quant, Context Size, and Token Counts? If any are missing, the screenshot is non-compliant [S2].

**Expanded Example Report Language**

**Scenario A: Standard Public Model Comparison**
"Model Comparison Report: [Model Name]
Configuration: [Quant] | Context: [Context Size]
Performance Metrics:
- Pass Rate: [X]%
- Invalid-run Count: [Y]
- Median Generation TPS: [Z]
- MTP Acceptance: [A]%
- Reasoning Tokens: [B]
- Final Tokens: [C]
- Output Artifacts: [Description]

*Privacy Note: All prompts used in this comparison have been redacted into summaries to comply with lab privacy standards [S6].*

**Scenario B: Internal Reliability Log (Failed Run)**
"Internal Log - Run ID: [ID]
Status: INVALID
Model: [Model Name]
Observation: The model failed to produce a valid output during the [Task Name] phase.
Reliability Analysis: [Description of the reliability problem identified].
Retention Status: Retained for reliability analysis per policy [S8].
Data Source: [WorkDash-derived artifact - INTERNAL ONLY] [S3]."

**Scenario C: Synthetic Prompt Export Log**
"Synthetic Prompt Export: [Prompt ID]
Verification Status: PASSED
- Real Names: None detected [S4]
- Emails: None detected [S4]
- Teams Messages: None detected [S4]
- Secrets: None detected [S4]
Export Authorized: Yes [S4]."

**Policy Interpretation and Conflict Resolution Log**

To ensure clarity for all lab members, the following logic was used to resolve conflicts between the synthetic sources:

1. **The "Retention" Conflict (S1 vs. S6):**
   - *Source S1 (2024-11):* Suggested a 14-day retention for all benchmark artifacts.
   - *Source S6 (2026-05):* Mandates that raw private prompts be retained until explicitly deleted.
   - *Resolution:* S6 is newer and more specific to "raw private prompts." Therefore, the "retain until deleted" rule is the active policy for prompts, while the 14-day rule [S1] is superseded for these specific items.

2. **The "Failed Run" Conflict (S7 vs. S8):**
   - *Source S7 (2025-08):* Suggested that all failed runs should be discarded.
   - *Source S8 (2026-06):* Mandates that failed and invalid runs should be retained and clearly labeled.
   - *Resolution:* S8 is the most recent policy (2026-06) and explicitly addresses the need to identify reliability problems. S8 takes precedence over S7.

3. **The "Reporting Style" Conflict (S2 vs. S6):**
   - *Source S2 (2025-05):* Focuses on the metadata requirements for screenshots (omitting prompts, including model/quant/context/tokens).
   - *Source S6 (2026-05):* Focuses on the general requirement for redacted summaries in publishable reports.
   - *Resolution:* These are complementary but S6 is newer. S6 provides the overarching rule for report content (redacted summaries), while S2 provides the specific metadata requirements for the visual elements (screenshots) of those reports.

**Compliance Checklist for Publication**

Before clicking "Publish," the following checklist must be completed:

- [ ] **WorkDash Check:** Is this artifact derived from WorkDash? (If yes, STOP - do not publish [S3]).
- [ ] **Prompt Check:** Are there any raw prompts? (If yes, replace with redacted summary [S6]).
- [ ] **Screenshot Metadata:** Does the screenshot show Model Name, Quant, Context Size, and Token Counts? (If no, add metadata [S2]).
- [ ] **Screenshot Prompt:** Is the raw prompt hidden in the screenshot? (If visible, re-take screenshot [S2]).
- [ ] **Metric Check:** Are all 7 metrics from S5 included? (Pass Rate, Invalid-run Count, Median TPS, MTP Acceptance, Reasoning Tokens, Final Tokens, Output Artifacts).
- [ ] **PII/Secret Scrub:** Have I verified there are no names, emails, Teams messages, or secrets? (If any exist, scrub them [S4]).

**Data Lifecycle Summary**

1. **Creation:** Model runs generate raw prompts and output artifacts.
2. **Classification:** Artifacts are tagged as "WorkDash-derived" [S3], "Raw Private" [S6], or "Synthetic" [S4].
3. **Processing:** Failed runs are labeled and retained [S8]. Raw prompts are kept locally [S6].
4. **Redaction:** Public reports are prepared by summarizing raw prompts [S6] and scrubbing synthetic prompts of PII/Secrets [S4].
5. **Metric Integration:** All model comparisons are populated with the 7 mandatory metrics [S5].
6. **Publication:** Screenshots are verified for metadata [S2] and the final report is released.

**Final Policy Statement**
The AI Flight Recorder home lab prioritizes the integrity of the research and the security of the data. By adhering to these policies, we ensure that our model comparisons are scientifically rigorous (via S5 and S8), our privacy is protected (via S3, S4, and S6), and our public documentation is clear and informative (via S2). All lab members are responsible for ensuring that no data leaves the lab in a non-compliant state.

# Source Priority
The policy resolution follows a strict hierarchy:
1. **Recency (Newest Policy):** Policies from 2026 (S3, S4, S5, S6, S8) take precedence over those from 2025 (S2, S7) and 2024 (S1).
2. **Specificity:** Where dates are equal or where a policy addresses a specific subset of data (e.g., "raw private prompts" vs. "benchmark artifacts"), the more specific policy is applied.
3. **Conflict Resolution:** 
    - S8 (2026-06) overrides S7 (2025-08) regarding failed runs.
    - S6 (2026-05) overrides S1 (2024-11) regarding prompt retention.
    - S6 (2026-05) overrides S2 (2025-05) regarding the use of redacted summaries in reports.

# Resolved Policy
**I. Artifact Retention Protocol**
- **Raw Private Prompts:** Must be stored locally and retained indefinitely until an explicit deletion command is issued [S6].
- **Failed/Invalid Runs:** Must be preserved and clearly tagged with a "Failure" or "Invalid" label to facilitate reliability debugging [S8].
- **General Benchmark Artifacts:** While a 14-day window was previously suggested [S1], the current mandate for raw prompt retention [S6] and failure retention [S8] establishes a "retain until deleted" standard for primary data.

**II. Publication and Export Rules**
- **Internal Only:** Any artifact derived from WorkDash is strictly confined to the home lab [S3].
- **Public Screenshots:** Must contain: Model Name, Quantization, Context Size, and Token Counts. They must *not* contain raw prompts [S2].
- **Public Reports:** Must use redacted summaries of prompts rather than the raw input [S6].
- **Synthetic Prompt Export:** Permitted only if the prompt is verified to contain zero real names, emails, Teams messages, or secrets [S4].

**III. Comparative Analysis Requirements**
Every model comparison report must include a standardized metrics block containing:
- Pass Rate [S5].
- Invalid-run Count [S5].
- Median Generation TPS [S5].
- MTP Acceptance [S5].
- Reasoning Tokens [S5].
- Final Tokens [S5].
- Output Artifacts [S5].

# Contradictions
The following contradictions were identified and resolved during the policy construction:
1. **Retention of Failed Runs:** S7 (2025-08) stated that failed runs should be discarded. This was overruled by S8 (2026-06), which mandates their retention and labeling to identify reliability issues. S8 is newer and therefore takes precedence.
2. **Retention Duration:** S1 (2024-11) suggested a 14-day retention for all benchmark artifacts. This was overruled by S6 (2026-05), which specifies that raw private prompts must be retained until explicitly deleted. S6 is newer and more specific.
3. **Reporting Style:** S2 (2025-05) focuses on the content of screenshots (omitting prompts but including metadata). S6 (2026-05) provides a broader reporting standard requiring redacted summaries for publishable reports. S6 is newer and takes precedence for the general reporting structure.

# Metrics To Report
The following metrics are mandatory for all model comparisons [S5]:
1. **Pass Rate:** The percentage of successful completions.
2. **Invalid-run Count:** The total number of runs that failed or produced invalid outputs.
3. **Median Generation TPS:** The median speed of token generation.
4. **MTP Acceptance:** The acceptance rate of Maximum Token Probability.
5. **Reasoning Tokens:** The count of tokens generated during the reasoning phase.
6. **Final Tokens:** The count of tokens in the final output.
7. **Output Artifacts:** The actual files or data structures produced by the model.

# What Must Stay Private
The following items are strictly prohibited from being published outside the home lab:
- **WorkDash-derived Artifacts:** Any data or artifacts originating from WorkDash [S3].
- **Raw Private Prompts:** These must remain local; only redacted summaries are allowed in public reports [S6].
- **Personally Identifiable Information (PII):** Real names and emails [S4].
- **Internal Communications:** Teams messages [S4].
- **Security Credentials:** Any secrets [S4].

# Example Report Language
**Example 1: Model Comparison Summary (Public)**
"Model Comparison: [Model Name] ([Quant])
Context Size: [Context Size]
Metrics:
- Pass Rate: [X]%
- Invalid-run Count: [Y]
- Median Generation TPS: [Z]
- MTP Acceptance: [A]%
- Reasoning Tokens: [B]
- Final Tokens: [C]
- Output Artifacts: [Description of artifacts]

*Note: Prompt details have been redacted for privacy [S6].*

**Example 2: Screenshot Caption (Public)**
"Screenshot of [Model Name] ([Quant]) at [Context Size] context. Total Token Count: [Count]. (Raw prompt omitted per policy [S2])."

**Example 3: Internal Failure Log (Private)**
"Run ID: [ID] - Status: INVALID/FAILED
Reason: [Description]
Note: Retained for reliability analysis [S8].
Source: WorkDash-derived artifact [S3] - DO NOT PUBLISH."

# Confidence
High. The policy was constructed by strictly adhering to the provided sources, resolving all chronological and specificity conflicts as instructed. All metrics and privacy constraints are directly mapped to the provided synthetic sources.