Compare Classification Thresholds Without Hiding Error-Cost Tradeoffs — Free
Two thresholds can trade false positives for false negatives while their accuracy percentages look reassuringly similar. Use a fixed classroom dataset and an explicit cost assumption to inspect that tradeoff instead of calling one threshold best without a criterion.
The proof surface
Practice-question coverage tools examine syllabus breadth. This board teaches confusion-matrix reconciliation: equal actual-class denominators, observed sensitivity and specificity, and a declared false-negative cost. It never deploys a model or interprets a medical test.
InputThreshold code | true positives | false positives | false negatives | true negatives; false-negative cost in false-positive cost units
Rare devicePractice-question coverage tools examine syllabus breadth. This board teaches confusion-matrix reconciliation: equal actual-class denominators, observed sensitivity and specificity, and a declared false-negative cost. It never deploys a model or interprets a medical test.
Output artifactLowest observed cost under the supplied error model with row-by-row context
Cost$0 local calculation · no card, paid key, subscription or signup · proposed filing price is not for sale
Sample, not your facts: Illustrative inputs: Conservative threshold | 8 | 2 | 2 | 8; Sensitive threshold | 9 | 6 | 1 | 4. False-negative cost in false-positive cost units = 5. Lowest observed cost under the supplied error model: 55 cost units per 100 records. All records are invented.
Before using the confusion-cost board
Every row must evaluate the same labeled records under a different threshold. True positives plus false negatives must therefore be identical across rows, as must false positives plus true negatives. Both actual classes must be present. These checks cannot verify the labels or prevent training-data leakage; they only reconcile counts. A false positive costs one model unit and a false negative costs the supplied multiple. True outcomes carry zero model cost. Costs may be zero, hypothetical or unsuitable for a real application. The summary selects minimum cost per hundred records and exposes tied minima without claiming a universal optimum. This educational exercise is not for diagnosing patients, assigning rights or automatically deciding consequential outcomes.
Why the flat version breaks
Comparing thresholds on different actual-class totals
Twenty records can contain ten positives in one row and only two in another. A count-only cost comparison then confounds threshold behavior with class mix. This parser requires both actual-class denominators to match, keeping the intended fixed-set experiment explicit rather than trusting equal overall totals.
Calling weighted cost an error percentage
With a false-negative multiplier of five, two errors can cost ten units. Cost per hundred can exceed one hundred and is not a percentage of wrong labels. The board uses cost-unit wording and displays the raw counts so a persuasive-looking number cannot be mistaken for accuracy.
Deploying the lowest training cost
Observed counts do not establish generalization, calibration, fairness or acceptable consequences. A training matrix may be perfect through leakage or memorization. Keep this educational comparison separate from real decisions and obtain a qualified evaluation of the dataset and deployment context.
How to work the confusion-cost board
Define the positive class and fixed evaluation set
Write what positive means and use one de-identified, non-sensitive classroom dataset for all thresholds. Keep the label source and split convention in the exercise notes. Switching datasets between rows changes the denominator and prevents a controlled threshold comparison, even if both rows happen to have the same total record count.
Reconcile the four observed counts
Enter whole true-positive, false-positive, false-negative and true-negative counts. Check both actual-class totals, not just the overall total. Each class needs a positive count for the displayed sensitivity and specificity. Missing labels cannot be treated as true negatives, and the app does not adjudicate whether a classification was genuinely correct.
Calculate the declared cost and its scale
For each threshold, compute false positives plus the supplied multiple times false negatives. Divide by total records and multiply by one hundred for comparable cost units per hundred. Also retain sensitivity and specificity so the same cost does not hide different error patterns. The model uses no guessed financial amount or claim about the severity of an error.
Discuss how the result changes with assumptions
Repeat the exercise with a justified different cost multiple and note which thresholds tie or change order. Ask an instructor or qualified analyst about evaluation design and real-world error costs. The free result remains a complete worked ledger; neither a cheap observed cost nor perfect training counts justify deployment.
What the confusion-cost board separates
Question
Before
Inspect this instead
Comparing thresholds on different actual-class totals
Twenty records can contain ten positives in one row and only two in another. A count-only cost comparison then confounds threshold behavior with class mix. This parser requires both actual-class denominators to match, keeping the intended fixed-set experiment explicit rather than trusting equal overall totals.
Reconcile the four observed counts
Calling weighted cost an error percentage
With a false-negative multiplier of five, two errors can cost ten units. Cost per hundred can exceed one hundred and is not a percentage of wrong labels. The board uses cost-unit wording and displays the raw counts so a persuasive-looking number cannot be mistaken for accuracy.
Calculate the declared cost and its scale
Deploying the lowest training cost
Observed counts do not establish generalization, calibration, fairness or acceptable consequences. A training matrix may be perfect through leakage or memorization. Keep this educational comparison separate from real decisions and obtain a qualified evaluation of the dataset and deployment context.
Discuss how the result changes with assumptions
This confusion-cost board replaces a manual count or calculation, not source verification or the responsible person’s review.
Run it on the samples, right here
FIRST-LOAD
HYPOTHESIS / PROTOTYPE — checkout unavailable. Calculation is local. A draft is saved automatically in this browser profile when storage is available; Reset to sample clears it. Optional Pro history stores only five summaries and has its own deletion control. State links encode your inputs and can remain in browser history, clipboard or recipients’ records; share only non-sensitive rows. Optional external AI formatting leaves this device. The required site analytics beacon reports page activity; shared URLs contain encoded inputs. Do not treat an encoded URL as private. The calculator has no input-collection endpoint.
Classroom evaluation arithmetic only, not medical-test interpretation, automated hiring or a deployment recommendation. Ask an instructor or qualified analyst to review labels, evaluation design and cost assumptions before any consequential use.
Data note: The confusion-cost board processes Threshold code | true positives | false positives | false negatives | true negatives locally. Starter/sample selection and Run compute in this tab; no input is sent by the calculator. A local draft may be saved; explicit state-link sharing or optional external AI formatting can disclose inputs. Use non-sensitive labels.
Go deeper: the companion app files the same reading as a triage clipboard with three lanes
The article demo above runs without limits. The companion app keeps a local history, exports the rows as CSV, prints the confusion-cost board reading, and holds your drafts on this device — one complete free app run; the proposed $4 one-time filing layer is not for sale.
The confusion-cost board answer stays complete for free. The proposed $4 one-time filing layer adds row CSV, print and five local reading summaries, not hidden answers. Checkout is unavailable; the article demo remains unlimited.
Classroom evaluation arithmetic only, not medical-test interpretation, automated hiring or a deployment recommendation. Ask an instructor or qualified analyst to review labels, evaluation design and cost assumptions before any consequential use.
What this is built on
Method: For each threshold, compute false positives plus the supplied multiple times false negatives. Divide by total records and multiply by one hundred for comparable cost units per hundred. Also retain sensitivity and specificity so the same cost does not hide different error patterns. The model uses no guessed financial amount or claim about the severity of an error.
All sample records, dates, quantities and labels are invented. No outside policy, contract, rate, clock offset, measurement or accessibility standard is represented as verified.
Google’s official pricing documentation, fetched 2026-10-01, says AI Studio is free in available regions. Optional formatting may require a Google account; manual local entry requires none. Limits can change and free-tier content may be used to improve products. Do not send private records.
Before: one accuracy number concealed unlike error types. After: a fixed dataset and explicit error-cost multiple make the threshold tradeoff inspectable.
Setting: False-negative cost in false-positive cost units = 5. Expected summary: 55 cost units per 100 records.
Both rows describe ten actual positive and ten actual negative records. Conservative has cost 2 + 5×2 = 12 over twenty records, or sixty cost units per hundred. Sensitive has cost 6 + 5×1 = 11, or fifty-five per hundred, despite more false positives. The summary reports the lowest modeled cost among the entered thresholds, not the highest accuracy or a deployment recommendation.
Setting: False-negative cost in false-positive cost units = 5. Expected summary: 65 cost units per 100 records.
Strict's cost is 1 + 5×4 = 21, or 105 units per hundred records. Balanced's cost is 3 + 5×2 = 13, or 65 units per hundred. A rate above one hundred is possible because each false negative is weighted at five cost units. It is not a percentage of incorrectly classified records and must not be labeled error percent.
Sample C — boundary convention
Perfect training threshold | 10 | 0 | 0 | 10
Setting: False-negative cost in false-positive cost units = 5. Expected summary: 0 cost units per 100 records.
The entered confusion matrix has no false positives or false negatives, so its observed cost is zero under this model. A perfect training result does not establish future performance or absence of bias. Both actual classes are present, providing defined sensitivity and specificity, but a new dataset or a different class prevalence can produce a different cost.
Optional AI formatting, never the calculation
Manual entry completes this confusion-cost board for free without signup. If available to you, the free AI Studio interface linked in the sources may format fictional or non-sensitive notes; external access may require an account. No API key or AI call is built into this tool. Free-tier content may be used to improve products. Review each cell and transcribe it to the labeled row schema; do not paste the JSON object into the row box.
Format only these fictional or non-sensitive notes for a confusion-cost board. Return strict JSON shaped as {"rows": [{"label": "string", "cells": ["string", "string", "string", "string"]}], "setting": "string"}. The columns are Threshold code | true positives | false positives | false negatives | true negatives; the setting is False-negative cost in false-positive cost units. Keep all supplied strings and quantities exactly; do not calculate, infer missing entries, invent dates or add advice. If any required value is missing, return an empty rows array and ask me for it separately. I will verify every cell against my source and manually transcribe rows using vertical bars before running the local calculator.
An AI response is not executed, fetched or trusted as a result. Missing values remain questions; the strict local parser checks the rows you actually enter.