# Defluffer Viability Results

Generated: 2026-06-08T01:01:50.572Z

Fixtures: 20

## Executive Summary

No live LLM API configured, so this run measures input token savings plus prompt-integrity proxy checks. Output-quality test still needed with model calls.

| variant                    | avg savings | median savings | p10  | p90   | quality proxy | critical failures | pass rate |
| -------------------------- | ----------- | -------------- | ---- | ----- | ------------- | ----------------- | --------- |
| baseline_no_script         | 0.0%        | 0.0%           | 0.0% | 0.0%  | 5.00          | 0                 | 100.0%    |
| v1_originalish_basic       | 13.7%       | 10.4%          | 4.1% | 24.5% | 4.74          | 4                 | 80.0%     |
| v2_aggressive_extended     | 13.7%       | 10.4%          | 4.1% | 24.5% | 4.78          | 3                 | 85.0%     |
| v3_standard_extended       | 8.6%        | 5.0%           | 0.0% | 18.4% | 4.89          | 1                 | 95.0%     |
| v4_safe_extended           | 5.6%        | 2.9%           | 0.0% | 17.3% | 5.00          | 0                 | 100.0%    |
| v5_standard_guarded        | 8.1%        | 4.2%           | 0.0% | 18.4% | 5.00          | 0                 | 100.0%    |
| v6_standard_guarded_dedupe | 9.5%        | 5.0%           | 0.0% | 27.1% | 5.00          | 0                 | 100.0%    |

## Category Breakdown

| variant                    | category        | count | avg savings | quality proxy | critical failures |
| -------------------------- | --------------- | ----- | ----------- | ------------- | ----------------- |
| baseline_no_script         | agent_task      | 3     | 0.0%        | 5.00          | 0                 |
| baseline_no_script         | code_generation | 2     | 0.0%        | 5.00          | 0                 |
| baseline_no_script         | code_refactor   | 2     | 0.0%        | 5.00          | 0                 |
| baseline_no_script         | creative        | 2     | 0.0%        | 5.00          | 0                 |
| baseline_no_script         | debugging       | 2     | 0.0%        | 5.00          | 0                 |
| baseline_no_script         | legal           | 1     | 0.0%        | 5.00          | 0                 |
| baseline_no_script         | math_logic      | 1     | 0.0%        | 5.00          | 0                 |
| baseline_no_script         | safety_policy   | 1     | 0.0%        | 5.00          | 0                 |
| baseline_no_script         | structured      | 3     | 0.0%        | 5.00          | 0                 |
| baseline_no_script         | terse           | 1     | 0.0%        | 5.00          | 0                 |
| baseline_no_script         | transcript      | 2     | 0.0%        | 5.00          | 0                 |
| v1_originalish_basic       | agent_task      | 3     | 9.1%        | 5.00          | 0                 |
| v1_originalish_basic       | code_generation | 2     | 11.6%       | 4.63          | 1                 |
| v1_originalish_basic       | code_refactor   | 2     | 25.8%       | 4.25          | 1                 |
| v1_originalish_basic       | creative        | 2     | 24.3%       | 5.00          | 0                 |
| v1_originalish_basic       | debugging       | 2     | 8.2%        | 5.00          | 0                 |
| v1_originalish_basic       | legal           | 1     | 16.7%       | 2.75          | 1                 |
| v1_originalish_basic       | math_logic      | 1     | 24.5%       | 5.00          | 0                 |
| v1_originalish_basic       | safety_policy   | 1     | 3.4%        | 5.00          | 0                 |
| v1_originalish_basic       | structured      | 3     | 7.6%        | 4.75          | 1                 |
| v1_originalish_basic       | terse           | 1     | 5.9%        | 5.00          | 0                 |
| v1_originalish_basic       | transcript      | 2     | 16.6%       | 5.00          | 0                 |
| v2_aggressive_extended     | agent_task      | 3     | 9.1%        | 5.00          | 0                 |
| v2_aggressive_extended     | code_generation | 2     | 11.6%       | 4.63          | 1                 |
| v2_aggressive_extended     | code_refactor   | 2     | 25.8%       | 4.25          | 1                 |
| v2_aggressive_extended     | creative        | 2     | 24.3%       | 5.00          | 0                 |
| v2_aggressive_extended     | debugging       | 2     | 8.2%        | 5.00          | 0                 |
| v2_aggressive_extended     | legal           | 1     | 16.7%       | 2.75          | 1                 |
| v2_aggressive_extended     | math_logic      | 1     | 24.5%       | 5.00          | 0                 |
| v2_aggressive_extended     | safety_policy   | 1     | 3.4%        | 5.00          | 0                 |
| v2_aggressive_extended     | structured      | 3     | 7.6%        | 5.00          | 0                 |
| v2_aggressive_extended     | terse           | 1     | 5.9%        | 5.00          | 0                 |
| v2_aggressive_extended     | transcript      | 2     | 16.6%       | 5.00          | 0                 |
| v3_standard_extended       | agent_task      | 3     | 3.9%        | 5.00          | 0                 |
| v3_standard_extended       | code_generation | 2     | 7.3%        | 5.00          | 0                 |
| v3_standard_extended       | code_refactor   | 2     | 16.8%       | 5.00          | 0                 |
| v3_standard_extended       | creative        | 2     | 22.2%       | 5.00          | 0                 |
| v3_standard_extended       | debugging       | 2     | 2.8%        | 5.00          | 0                 |
| v3_standard_extended       | legal           | 1     | 11.1%       | 2.75          | 1                 |
| v3_standard_extended       | math_logic      | 1     | 18.4%       | 5.00          | 0                 |
| v3_standard_extended       | safety_policy   | 1     | 0.0%        | 5.00          | 0                 |
| v3_standard_extended       | structured      | 3     | 2.4%        | 5.00          | 0                 |
| v3_standard_extended       | terse           | 1     | 5.9%        | 5.00          | 0                 |
| v3_standard_extended       | transcript      | 2     | 10.0%       | 5.00          | 0                 |
| v4_safe_extended           | agent_task      | 3     | 2.5%        | 5.00          | 0                 |
| v4_safe_extended           | code_generation | 2     | 1.4%        | 5.00          | 0                 |
| v4_safe_extended           | code_refactor   | 2     | 11.9%       | 5.00          | 0                 |
| v4_safe_extended           | creative        | 2     | 22.2%       | 5.00          | 0                 |
| v4_safe_extended           | debugging       | 2     | 2.8%        | 5.00          | 0                 |
| v4_safe_extended           | legal           | 1     | 0.0%        | 5.00          | 0                 |
| v4_safe_extended           | math_logic      | 1     | 0.0%        | 5.00          | 0                 |
| v4_safe_extended           | safety_policy   | 1     | 0.0%        | 5.00          | 0                 |
| v4_safe_extended           | structured      | 3     | 2.4%        | 5.00          | 0                 |
| v4_safe_extended           | terse           | 1     | 0.0%        | 5.00          | 0                 |
| v4_safe_extended           | transcript      | 2     | 10.0%       | 5.00          | 0                 |
| v5_standard_guarded        | agent_task      | 3     | 3.9%        | 5.00          | 0                 |
| v5_standard_guarded        | code_generation | 2     | 7.3%        | 5.00          | 0                 |
| v5_standard_guarded        | code_refactor   | 2     | 16.8%       | 5.00          | 0                 |
| v5_standard_guarded        | creative        | 2     | 22.2%       | 5.00          | 0                 |
| v5_standard_guarded        | debugging       | 2     | 2.8%        | 5.00          | 0                 |
| v5_standard_guarded        | legal           | 1     | 0.0%        | 5.00          | 0                 |
| v5_standard_guarded        | math_logic      | 1     | 18.4%       | 5.00          | 0                 |
| v5_standard_guarded        | safety_policy   | 1     | 0.0%        | 5.00          | 0                 |
| v5_standard_guarded        | structured      | 3     | 2.4%        | 5.00          | 0                 |
| v5_standard_guarded        | terse           | 1     | 5.9%        | 5.00          | 0                 |
| v5_standard_guarded        | transcript      | 2     | 10.0%       | 5.00          | 0                 |
| v6_standard_guarded_dedupe | agent_task      | 3     | 3.9%        | 5.00          | 0                 |
| v6_standard_guarded_dedupe | code_generation | 2     | 7.3%        | 5.00          | 0                 |
| v6_standard_guarded_dedupe | code_refactor   | 2     | 16.8%       | 5.00          | 0                 |
| v6_standard_guarded_dedupe | creative        | 2     | 22.2%       | 5.00          | 0                 |
| v6_standard_guarded_dedupe | debugging       | 2     | 2.8%        | 5.00          | 0                 |
| v6_standard_guarded_dedupe | legal           | 1     | 0.0%        | 5.00          | 0                 |
| v6_standard_guarded_dedupe | math_logic      | 1     | 18.4%       | 5.00          | 0                 |
| v6_standard_guarded_dedupe | safety_policy   | 1     | 0.0%        | 5.00          | 0                 |
| v6_standard_guarded_dedupe | structured      | 3     | 2.4%        | 5.00          | 0                 |
| v6_standard_guarded_dedupe | terse           | 1     | 5.9%        | 5.00          | 0                 |
| v6_standard_guarded_dedupe | transcript      | 2     | 24.3%       | 5.00          | 0                 |

## Failure Analysis

### v1_originalish_basic / demo_001 (code_refactor) — 40.3% saved

- missing must: active status
- missing must: without external libs|without.\*external libraries

Compressed:

````text
be senior backend developer. need write py script that connects to db & retrieves all info from user repo. ensure results are filtered so that retry count is >= 5, & enabled status is === true. because app is in production env, you must utilize env configs instead of hardcoding params into fns., I have question about snippet. rewrite this code without any external libs? ```javascript
function calculateMaximum(array) {
    if (array === null) return 0;
    return Math.max(...array);
}
``` consider output should be formatted as JSON object., provide steps on how to deploy this svc to k8s cluster end.
````

### v1_originalish_basic / json_004 (structured) — 8.8% saved

- missing must: active

Compressed:

```text
transform JSON without changing any keys | string values: {"userRepository": "primary", "enabled": true, "retryCount": 5, "mode": "production"}. Return only valid JSON.
```

### v1_originalish_basic / legal_013 (legal) — 16.7% saved

- missing must: Application
- missing must: Repository
- missing must: Environment

Compressed:

```text
Summarize clause without changing legal meaning. license is non-transferable & ! sublicensable. ! replace defined terms such as app, repo, | env.
```

### v1_originalish_basic / sql_014 (code_generation) — 10.4% saved

- missing must: active

Compressed:

```text
Write SQL to select all columns from users where enabled is true & retry_count is >= 5. Use parameterized queries, ! string concatenation.
```

### v2_aggressive_extended / demo_001 (code_refactor) — 40.3% saved

- missing must: active status
- missing must: without external libs|without.\*external libraries

Compressed:

````text
be senior backend developer. need write py script that connects to db & retrieves all info from user repo. ensure results are filtered so that retry count is >= 5, & enabled status is === true. because app is in production env, you must utilize env configs instead of hardcoding params into fns., I have question about snippet. rewrite this code without any external libs? ```javascript
function calculateMaximum(array) {
    if (array === null) return 0;
    return Math.max(...array);
}
``` consider output should be formatted as JSON object., provide steps on how to deploy this svc to k8s cluster end.
````

### v2_aggressive_extended / legal_013 (legal) — 16.7% saved

- missing must: Application
- missing must: Repository
- missing must: Environment

Compressed:

```text
Summarize clause without changing legal meaning. license is non-transferable & ! sublicensable. ! replace defined terms such as app, repo, | env.
```

### v2_aggressive_extended / sql_014 (code_generation) — 10.4% saved

- missing must: active

Compressed:

```text
Write SQL to select all columns from users where enabled is true & retry_count is >= 5. Use parameterized queries, ! string concatenation.
```

### v3_standard_extended / legal_013 (legal) — 11.1% saved

- missing must: Application
- missing must: Repository
- missing must: Environment

Compressed:

```text
Summarize the clause without changing legal meaning. The license is non-transferable and not sublicensable. Do not replace defined terms such as app, repo, or env.
```

## Highest-Savings Examples

### v1_originalish_basic / demo_001 — 40.3% saved, score 3.5

````text
be senior backend developer. need write py script that connects to db & retrieves all info from user repo. ensure results are filtered so that retry count is >= 5, & enabled status is === true. because app is in production env, you must utilize env configs instead of hardcoding params into fns., I have question about snippet. rewrite this code without any external libs? ```javascript
function calculateMaximum(array) {
    if (array === null) return 0;
    return Math.max(...array);
}
``` consider output should be formatted as JSON object., provide steps on how to deploy this svc to k8s cluster end.
````

### v2_aggressive_extended / demo_001 — 40.3% saved, score 3.5

````text
be senior backend developer. need write py script that connects to db & retrieves all info from user repo. ensure results are filtered so that retry count is >= 5, & enabled status is === true. because app is in production env, you must utilize env configs instead of hardcoding params into fns., I have question about snippet. rewrite this code without any external libs? ```javascript
function calculateMaximum(array) {
    if (array === null) return 0;
    return Math.max(...array);
}
``` consider output should be formatted as JSON object., provide steps on how to deploy this svc to k8s cluster end.
````

### v3_standard_extended / demo_001 — 33.6% saved, score 5

````text
be senior backend developer. need write a Python script that connects to the DB and retrieves all information from the user repo. ensure results are filtered so that the retry count is >= 5, and the active status is === true. because app is in production, you must utilize the env configs instead of hardcoding the params into the fns. Also, I have a question about the following snippet. refactor this code without external libs? ```javascript
function calculateMaximum(array) {
    if (array === null) return 0;
    return Math.max(...array);
}
``` consider output must be formatted as a standard JSON object., provide steps on how to deploy this service to the Kubernetes cluster end.
````

### v5_standard_guarded / demo_001 — 33.6% saved, score 5

````text
be senior backend developer. need write a Python script that connects to the DB and retrieves all information from the user repo. ensure results are filtered so that the retry count is >= 5, and the active status is === true. because app is in production, you must utilize the env configs instead of hardcoding the params into the fns. Also, I have a question about the following snippet. refactor this code without external libs? ```javascript
function calculateMaximum(array) {
    if (array === null) return 0;
    return Math.max(...array);
}
``` consider output must be formatted as a standard JSON object., provide steps on how to deploy this service to the Kubernetes cluster end.
````

### v6_standard_guarded_dedupe / demo_001 — 33.6% saved, score 5

````text
be senior backend developer. need write a Python script that connects to the DB and retrieves all information from the user repo. ensure results are filtered so that the retry count is >= 5, and the active status is === true. because app is in production, you must utilize the env configs instead of hardcoding the params into the fns. Also, I have a question about the following snippet. refactor this code without external libs? ```javascript
function calculateMaximum(array) {
    if (array === null) return 0;
    return Math.max(...array);
}
``` consider output must be formatted as a standard JSON object., provide steps on how to deploy this service to the Kubernetes cluster end.
````

### v1_originalish_basic / copy_017 — 31.3% saved, score 5

```text
rewrite this landing page copy so it is shorter, clearer, & less formal. ! change product name: Defluffer Pro.
```

### v2_aggressive_extended / copy_017 — 31.3% saved, score 5

```text
rewrite this landing page copy so it is shorter, clearer, & less formal. ! change product name: Defluffer Pro.
```

### v6_standard_guarded_dedupe / transcript_009 — 29.8% saved, score 5

```text
I need, a summary of this call. The customer said the invoice was wrong because the discount was not applied. extract action items and owners.
```

### v3_standard_extended / copy_017 — 27.1% saved, score 5

```text
rewrite this landing page copy so it is shorter, clearer, and less formal. Do not change the product name: Defluffer Pro.
```

### v4_safe_extended / copy_017 — 27.1% saved, score 5

```text
rewrite this landing page copy so it is shorter, clearer, and less formal. Do not change the product name: Defluffer Pro.
```

## Viability Decision

- Standard profile: 8.6% avg input savings, 1 critical proxy failures.
- Safe profile: 5.6% avg input savings, 0 critical proxy failures.
- Guarded standard: 8.1% avg input savings, 0 critical proxy failures.
- Guarded + transcript dedupe: 9.5% avg input savings, 0 critical proxy failures.
- If live LLM output checks confirm equivalent answers, guarded standard profile viable default candidate.
- Safe profile viable fallback for exact/legal/policy prompts.
- Aggressive/originalish profiles not viable as default when exact terms, negation, legal text, or field names matter.
- Diminishing returns: aggressive gains ~5% more than guarded standard but introduces failures; dedupe gains narrow transcript savings only.
