TrustQueryNet — trustworthy classification under label noise
A 5-seed study on medical images where my own matched-budget control beat the method — and I published the table.
Group-aware lesion-level splits, per-seed CSVs, resolved config per run, and paired significance tests — the best experimental hygiene in the portfolio.
Results
Every value links to the committed artifact that produces it.
| Metric | Value | Evidence |
|---|---|---|
| HAM10000 accuracy (5 seeds)random-repair control: 0.8319 ± 0.0051 — statistically a tie | 0.8350 ± 0.0059 | artifacts/paper_tables/noisy_anchor/ablation_table.json |
| Calibration error (ECE) | 0.0445 ± 0.0090 | artifacts/paper_tables/noisy_anchor/ablation_table.json |
| External shift, ISIC-2019 accuracyECE rises to 0.20 — distribution shift is the real obstacle, and it was measured | 0.5692 ± 0.0145 | artifacts/paper_tables/external_main/ablation_table.csv |
| Duplicate images across splits (overlap audit) | 0 | artifacts/paper_tables/significance/paired_significance.csv |
What didn’t work
The matched-budget random-repair control matched the method within noise. The manuscript states plainly that the internal comparison does not support a "repair wins everywhere" story.
Reported as a null, not hidden — the same honesty policy applies to every number on this site.
Running it
No live demo by design: a skin-image endpoint open to the internet would hand strangers a silent "medical" answer. The committed reliability diagram, risk–coverage curves and ablation table stand in for it.