fix(recipes): recalibre CONFIDENCE_THRESHOLD pour le nouveau classifieur spaCy

0.75 (calibre a l'origine contre node-nlp) laissait de vrais verdicts
corrects sur des clauses sans ancre NER (rien sur quoi retomber) sous le
seuil : melt scorait 0.68 sur "jusqu'a ce que le beurre ait disparu dans
la poele" (le cas motivant tout ce pipeline), preheat 0.52 sur "mettre la
poele sur feu vif" — tous deux corrects, tous deux rejetes a 0.75.

Recalibre a 0.45 : marge confortable au-dessus du bruit (texte anglais
via le classifieur francais score ~0.04, indiscernable du hasard sur ~26
classes) et sous les deux cas ci-dessus. Confirme par
calibrate-tech-step-threshold.ts contre TECH_STEP_EVAL_DATASET (F1
plafonne a 0.987 des 0.45, reste plat jusqu'a 0.95 — 0.45 est deja le
seuil le plus bas qui capture tout le gain disponible).

Retire l'etape CI temporaire de calibration (ci.yml) une fois la valeur
choisie.

Verifie : lint + build complets du monorepo, 27/27 pytest du service,
sweep de seuils + verification manuelle contre le corpus reel en local
(services Python, sans Postgres) et en CI. La suite Mocha complete reste
a confirmer sur ce commit (executee en CI, pas localement — pas de
Postgres disponible dans cet environnement).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
Nicolas 2026-08-25 20:43:41 +02:00
parent 6128400414
commit 9590569f3f
2 changed files with 24 additions and 16 deletions

View file

@ -98,11 +98,6 @@ jobs:
- run: pnpm install --frozen-lockfile
- run: pnpm --filter api exec prisma migrate deploy
# TEMPORAIRE — recalibration de CONFIDENCE_THRESHOLD (tech-step-matcher.ts)
# contre le vrai service Python démarré ci-dessus. À retirer une fois
# la valeur choisie et reportée dans le code (voir le plan de migration).
- run: pnpm --filter api exec prisma db seed
- run: pnpm --filter api exec tsx src/scripts/calibrate-tech-step-threshold.ts
- run: pnpm --filter api test
intent-service-test:

View file

@ -242,18 +242,31 @@ export function splitIntoClauses(
* `TECH_STEP_TRAINING_DATA` see `test/tech-step-matcher.test.ts` for the
* cases this threshold was picked to pass.
*
* Raised from `0.65` after finding real (non-adversarial) misclassified
* clauses that scored just above the old threshold e.g. English recipe
* text run through the French classifier (which must find *nothing*,
* confirmed by `recipe-translation.test.ts`'s own locale-isolation test)
* scored `0.69` for `boil`, essentially classifier noise on
* out-of-vocabulary input rather than a real, confident verdict. The
* clauses this threshold exists to actually trust score far higher in
* practice (`0.91``1.0` for the real corrected cases found this session)
* `0.75` sits comfortably above the noise floor and below every genuine
* match seen so far.
* Recalibrated to `0.45` for the migration off `node-nlp` to
* `services/tech-step-intent-service` (spaCy `textcat`, exclusive classes
* over ~26 techniques) its score distribution is meaningfully different
* from node-nlp's own classifier, and the old `0.75` (tuned against
* node-nlp) left genuine, correct verdicts on *anchor-less* clauses
* (nothing to fall back to see this file's doc comment, point 3) below
* the bar: e.g. `melt` scored `0.68` on "jusqu'à ce que le beurre ait
* disparu dans la poêle" (the exact motivating no-keyword case this
* pipeline exists for), `preheat` scored `0.52` on "mettre la poêle sur
* feu vif" both the classifier's own confident best guess, both correct,
* both wrongly discarded at `0.75`. `0.45` clears both with margin while
* staying far above the noise floor: English recipe text run through the
* French classifier (must find *nothing*, confirmed by
* `recipe-translation.test.ts`'s own locale-isolation test) scored `0.04`
* for every technique indistinguishable from the ~1/26 uniform baseline
* over this many exclusive classes, nowhere near `0.45`. Cross-checked
* against `apps/api/src/scripts/calibrate-tech-step-threshold.ts`'s sweep
* over `TECH_STEP_EVAL_DATASET`: aggregate F1 climbs to its plateau
* (`0.987`) starting exactly at `0.45` and stays flat through `0.95`, so
* this is the lowest threshold that already captures every gain available
* from trusting the classifier more a higher value would only ever
* discard more anchor-less true positives like the two above, never buy
* back any precision.
*/
export const CONFIDENCE_THRESHOLD = 0.75;
export const CONFIDENCE_THRESHOLD = 0.45;
/**
* One clause's full classification detail the finer-grained sibling of