* feat(recipes): ajoute un adaptateur RecipeSourceAdapter pour Marmiton Étend jsonLdRecipeAdapter (json-ld-recipe.ts) plutôt que de dupliquer sa logique : marmitonAdapter délègue fetchDetail/parse directement à l'adaptateur générique JSON-LD (une page recette marmiton.org expose un Recipe schema.org standard), et n'ajoute que ce que l'adaptateur générique ne peut pas offrir — un list() qui lit l'ItemList schema.org embarqué sur la page de résultats de recherche de marmiton.org (pagination via &page=N, fin de résultats détectée via la réponse 404 renvoyée au-delà de la dernière page). extractJsonLdBlocks est exporté depuis json-ld-recipe.ts pour être réutilisé par marmiton.ts sans dupliquer le regex d'extraction des blocs <script type="application/ld+json">. Enregistre marmitonAdapter dans registerAllRecipeSources (sources/index.ts) — contrairement à jsonLdRecipeAdapter lui-même, c'est un adaptateur concret par site, donc une Source household-toggleable légitime. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * feat(recipes): ajoute un adaptateur RecipeSourceAdapter pour 750g Suit le même schéma que marmitonAdapter (construit sur jsonLdRecipeAdapter), avec deux différences propres à 750g.com : - list() n'a pas d'ItemList JSON-LD à lire sur ses résultats de recherche (la recherche du site est un widget client-side) — appelle donc directement le endpoint GET que ce widget interroge lui-même en interne (un « moteur de réponse IA » qui renvoie un lot de recettes pour une requête en texte libre), et scrape les cartes de résultat par regex en associant à chaque lien de recette sa dernière image précédente plutôt qu'un zip naïf par index (des images décoratives sans carte associée existent réellement dans ce fragment). Vérifié en direct : demander une « page 2 » revient toujours vide, donc nextCursor vaut toujours null, comme theMealDbAdapter. - parse() ne délègue pas aussi directement à jsonLdRecipeAdapter.parse que marmitonAdapter — le générateur JSON-LD de 750g.com a deux bugs réels : des caractères de contrôle bruts non échappés dans certaines chaînes JSON (~1 recette sur 3 dans un échantillon vérifié en direct, sinon JSON.parse échoue et jsonLdRecipeAdapter rapporte à tort « aucun Recipe trouvé »), et un texte parfois doublement encodé en entités HTML (ex. un vrai « é » devient &eacute; au lieu de é). Les deux sont corrigés en pré/post-traitement autour de la même délégation, pas une réimplémentation. Enregistre sevenFiftyGAdapter dans registerAllRecipeSources (sources/index.ts), au même titre que marmitonAdapter. Complète aussi test/sources/sources-index.test.ts, qui ne couvrait encore que TheMealDB malgré l'ajout de Marmiton dans une PR précédente. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * feat(recipes): ajoute un adaptateur RecipeSourceAdapter pour Manger Bouger Suit le même schéma que marmitonAdapter/sevenFiftyGAdapter (construit sur jsonLdRecipeAdapter), avec des différences propres à mangerbouger.fr (« La Fabrique à Menus », Santé publique France) : - list() n'utilise pas de JSON-LD du tout — la page de résultats (une app Next.js) n'embarque aucun ItemList. Elle est cependant rendue côté serveur et expose le même state Redux que le client hydrate, via un <script id="__NEXT_DATA__">, qui contient déjà tout ce dont list() a besoin (slug/nom/image, pagination). Vérifié en direct : ?query=<texte libre> filtre bien côté serveur, et hasMorePages donne un signal de fin de pagination plus propre que le 404 de Marmiton ou l'absence de vraie pagination de 750g. - parse() délègue à jsonLdRecipeAdapter mais corrige deux lacunes réelles et systématiques de son propre JSON-LD (vérifiées sur 9 recettes, 72 étapes) : recipeInstructions[].text est un document Slate.js sérialisé en JSON (pas du texte) plutôt qu'être aplati ; recipeYield est absent partout alors que le nombre de portions existe bien côté site (__NEXT_DATA__) — les deux sont corrigés par un patch structuré (parse → mutation → réécriture) avant délégation, pas une réimplémentation. Enregistre mangerBougerAdapter dans registerAllRecipeSources (sources/index.ts) et complète sources-index.test.ts. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
355 lines
15 KiB
TypeScript
355 lines
15 KiB
TypeScript
import type {
|
|
ParsedRecipe,
|
|
RecipeSourceAdapter,
|
|
RecipeSourceListItem,
|
|
RecipeSourceListParams,
|
|
RecipeSourceListResult,
|
|
} from "../lib/recipe-sources/recipe-source-adapter.js";
|
|
import {
|
|
RecipeSourceFetchError,
|
|
RecipeSourceParseError,
|
|
} from "../lib/recipe-sources/recipe-source-errors.js";
|
|
import { jsonLdRecipeAdapter } from "./json-ld-recipe.js";
|
|
|
|
const SOURCE_KEY = "mangerBouger";
|
|
|
|
// "La Fabrique à Menus" — mangerbouger.fr's recipe tool (Santé publique
|
|
// France). Its listing page is a Next.js app with no JSON-LD `ItemList` at
|
|
// all (unlike marmiton.ts's search page) — but it's server-rendered, and a
|
|
// plain GET carries the exact same Redux state the client hydrates from as
|
|
// a `__NEXT_DATA__` script tag (see `extractNextData` below), which already
|
|
// has everything `list()` needs. Verified live: `?query=<free text>` really
|
|
// filters server-side (not just a client-side URL update over an
|
|
// already-fetched page), and `page`/`hasMorePages` behave as real,
|
|
// consistent pagination — the best-behaved of this adapter family's three
|
|
// sources on that front.
|
|
const LIST_URL = "https://www.mangerbouger.fr/manger-mieux/la-fabrique-a-menus/recettes";
|
|
const DETAIL_BASE_URL = "https://www.mangerbouger.fr/manger-mieux/la-fabrique-a-menus/recettes/";
|
|
|
|
/** Matches the `<script id="__NEXT_DATA__">…</script>` block every Next.js page ships — the site's own server-rendered hydration data, read instead of scraping HTML for both `list()` (the listing's recipe cards) and `parse()` (backfilling a gap in the detail page's JSON-LD, see {@link extractPortionsFromNextData}). */
|
|
const NEXT_DATA_PATTERN = /<script id="__NEXT_DATA__"[^>]*>([\s\S]*?)<\/script>/;
|
|
|
|
/** The one field of one `list[]` entry `list()` actually reads off the listing page's `__NEXT_DATA__` — that state carries the site's full internal `Recipe` shape (60+ fields: nutriscore, seasons, macros, …), none of which this adapter's contract has anywhere to put. */
|
|
interface MangerBougerListEntry {
|
|
slug?: string;
|
|
name?: string;
|
|
image?: string | null;
|
|
}
|
|
|
|
/** The slice of `__NEXT_DATA__` this module reads off the *listing* page. */
|
|
interface MangerBougerListPageData {
|
|
props?: {
|
|
initialState?: {
|
|
recipes?: {
|
|
list?: MangerBougerListEntry[];
|
|
/** Whether a further page exists for the current `page`/`query`/`diet` combination — verified live: an out-of-range page comes back `false` with an empty `list` rather than repeating the last page or erroring, a cleaner end-of-results signal than either `marmiton.ts` (infers it from a 404) or `750g.ts` (this search has no real pagination at all). */
|
|
hasMorePages?: boolean;
|
|
};
|
|
};
|
|
};
|
|
}
|
|
|
|
/** The slice of `__NEXT_DATA__` this module reads off a recipe *detail* page — a different shape than the listing page's (`initialState.recipe.recipe`, not `initialState.recipes.list[]`) since it's a different Redux slice entirely. */
|
|
interface MangerBougerDetailPageData {
|
|
props?: {
|
|
initialState?: {
|
|
recipe?: {
|
|
recipe?: {
|
|
portions?: unknown;
|
|
};
|
|
};
|
|
};
|
|
};
|
|
}
|
|
|
|
/** Parses the page's `__NEXT_DATA__` block into `T`, or `null` if the block is missing or isn't valid JSON — callers degrade gracefully rather than throw, same as `marmiton.ts`'s "page has no ItemList at all" handling. */
|
|
function extractNextData<T>(html: string): T | null {
|
|
const match = html.match(NEXT_DATA_PATTERN);
|
|
if (!match) return null;
|
|
try {
|
|
return JSON.parse(match[1] ?? "") as T;
|
|
} catch {
|
|
return null;
|
|
}
|
|
}
|
|
|
|
function detailUrl(slug: string): string {
|
|
return `${DETAIL_BASE_URL}${slug}`;
|
|
}
|
|
|
|
/**
|
|
* Matches the single `<script type="application/ld+json">…</script>` block
|
|
* a mangerbouger.fr recipe *detail* page carries (verified live across a
|
|
* sample of 9 recipes — always exactly one, always a bare `Recipe`, never
|
|
* an `@graph`) — a much narrower pattern than `json-ld-recipe.ts`'s own
|
|
* `JSON_LD_SCRIPT_PATTERN` (no `g` flag: this module only ever needs the
|
|
* first/only block, to patch it — see {@link patchRecipeJsonLd}) or
|
|
* `750g.ts`'s identically-named private copy (which does its own,
|
|
* different, character-level repair over every block on the page).
|
|
*/
|
|
const JSON_LD_SCRIPT_PATTERN =
|
|
/(<script[^>]*type\s*=\s*["']application\/ld\+json["'][^>]*>)([\s\S]*?)(<\/script>)/i;
|
|
|
|
/** One node of a Slate.js rich-text document — see {@link flattenSlateDocument}. */
|
|
interface SlateNode {
|
|
type?: string;
|
|
text?: string;
|
|
children?: SlateNode[];
|
|
}
|
|
|
|
/** Concatenates a run of inline Slate nodes (leaf text, or further-nested inline runs) with no separator — bold/italic/underline marks (the only ones observed) carry no plain-text equivalent and are simply dropped. */
|
|
function flattenSlateInline(nodes: SlateNode[]): string {
|
|
return nodes
|
|
.map((node) =>
|
|
typeof node.text === "string"
|
|
? node.text
|
|
: node.children
|
|
? flattenSlateInline(node.children)
|
|
: "",
|
|
)
|
|
.join("");
|
|
}
|
|
|
|
/**
|
|
* Flattens a Slate.js document's top-level blocks into one line of plain
|
|
* text each — verified live across every recipe step sampled (72 recipes):
|
|
* only `paragraph` and `bulleted-list` (of `list-item`s) ever appear as
|
|
* block types, so that's all this handles; any other/unrecognized block
|
|
* type still degrades reasonably (its own children read as one inline run)
|
|
* rather than being dropped outright.
|
|
*/
|
|
function flattenSlateBlocks(nodes: SlateNode[]): string[] {
|
|
const lines: string[] = [];
|
|
for (const node of nodes) {
|
|
if (node.type === "bulleted-list" && node.children) {
|
|
lines.push(...flattenSlateBlocks(node.children));
|
|
continue;
|
|
}
|
|
if (node.type === "list-item" && node.children) {
|
|
const text = flattenSlateInline(node.children);
|
|
if (text.trim().length > 0) lines.push(`- ${text}`);
|
|
continue;
|
|
}
|
|
if (node.children) {
|
|
const text = flattenSlateInline(node.children);
|
|
if (text.trim().length > 0) lines.push(text);
|
|
continue;
|
|
}
|
|
if (typeof node.text === "string" && node.text.trim().length > 0) lines.push(node.text);
|
|
}
|
|
return lines;
|
|
}
|
|
|
|
/**
|
|
* Flattens one `HowToStep.text` value into plain text. mangerbouger.fr's
|
|
* own JSON-LD embeds this field pre-formatted for its own web app instead
|
|
* of as prose: `text` is itself a JSON-serialized Slate.js rich-text
|
|
* document (verified live: every one of 72 sampled recipe steps parses as
|
|
* one) — handing that straight to `jsonLdRecipeAdapter.parse` would surface
|
|
* the raw `[{"type":"paragraph","children":[{"text":"…` blob as a step's
|
|
* description, unusable as-is. `json` that doesn't parse as an array (a
|
|
* genuinely plain-text step, or some future/different shape) is returned
|
|
* unchanged rather than mangled.
|
|
*/
|
|
function flattenSlateDocument(json: string): string {
|
|
let doc: unknown;
|
|
try {
|
|
doc = JSON.parse(json);
|
|
} catch {
|
|
return json;
|
|
}
|
|
if (!Array.isArray(doc)) return json;
|
|
return flattenSlateBlocks(doc as SlateNode[]).join("\n");
|
|
}
|
|
|
|
/** The two schema.org `Recipe` fields {@link patchRecipeJsonLd} patches, plus an index signature so every other field survives re-serialization untouched. */
|
|
interface JsonLdRecipeLike {
|
|
recipeInstructions?: unknown;
|
|
recipeYield?: unknown;
|
|
[key: string]: unknown;
|
|
}
|
|
|
|
/** One `HowToStep`-shaped entry of `recipeInstructions`, as far as {@link patchRecipeJsonLd} needs to know. */
|
|
interface JsonLdHowToStepLike {
|
|
text?: unknown;
|
|
[key: string]: unknown;
|
|
}
|
|
|
|
/**
|
|
* `state.recipe.recipe.portions` from the same detail page's `__NEXT_DATA__`
|
|
* — the number `recipeYield` should have been (see {@link patchRecipeJsonLd}),
|
|
* read from the site's own internal state rather than left unstated.
|
|
*/
|
|
function extractPortionsFromNextData(html: string): number | null {
|
|
const data = extractNextData<MangerBougerDetailPageData>(html);
|
|
const portions = data?.props?.initialState?.recipe?.recipe?.portions;
|
|
return typeof portions === "number" ? portions : null;
|
|
}
|
|
|
|
/**
|
|
* Repairs the two real gaps verified live in mangerbouger.fr's own
|
|
* recipe-detail JSON-LD, then hands the patched HTML to
|
|
* `jsonLdRecipeAdapter.parse` unmodified otherwise — same "fix what's
|
|
* actually broken, delegate the rest" shape as `750g.ts`'s
|
|
* `sanitizeJsonLdBlocks`/`decodeParsedRecipeText`, just structural (parse →
|
|
* mutate → re-serialize the one JSON-LD object) rather than textual, since
|
|
* both gaps need real understanding of the document, not character-level
|
|
* fixups:
|
|
*
|
|
* - `recipeInstructions[].text` is Slate.js rich text, not prose — flattened
|
|
* via {@link flattenSlateDocument}.
|
|
* - `recipeYield` is absent on every one of 9 sampled recipes (schema.org
|
|
* allows omitting it, and mangerbouger.fr's generator apparently always
|
|
* does) even though the site's own internal data has the serving count
|
|
* right there — backfilled from `__NEXT_DATA__` via
|
|
* {@link extractPortionsFromNextData} rather than left as a needless
|
|
* `portions: null` on every single imported recipe.
|
|
*
|
|
* A missing or malformed JSON-LD block is left completely untouched —
|
|
* `jsonLdRecipeAdapter`'s own "no JSON-LD Recipe found"/"malformed block,
|
|
* skip it" handling is exactly the right behavior for that, no need to
|
|
* duplicate it here.
|
|
*/
|
|
function patchRecipeJsonLd(html: string): string {
|
|
const match = html.match(JSON_LD_SCRIPT_PATTERN);
|
|
if (!match) return html;
|
|
|
|
let recipe: JsonLdRecipeLike;
|
|
try {
|
|
recipe = JSON.parse(match[2] ?? "{}") as JsonLdRecipeLike;
|
|
} catch {
|
|
return html;
|
|
}
|
|
|
|
if (Array.isArray(recipe.recipeInstructions)) {
|
|
for (const step of recipe.recipeInstructions as JsonLdHowToStepLike[]) {
|
|
if (step && typeof step === "object" && typeof step.text === "string") {
|
|
step.text = flattenSlateDocument(step.text);
|
|
}
|
|
}
|
|
}
|
|
|
|
if (recipe.recipeYield === undefined) {
|
|
const portions = extractPortionsFromNextData(html);
|
|
if (portions !== null) recipe.recipeYield = portions;
|
|
}
|
|
|
|
const patchedJson = JSON.stringify(recipe);
|
|
return html.replace(
|
|
JSON_LD_SCRIPT_PATTERN,
|
|
(_full, openTag: string, _json: string, closeTag: string) =>
|
|
`${openTag}${patchedJson}${closeTag}`,
|
|
);
|
|
}
|
|
|
|
/**
|
|
* Re-labels a `RecipeSourceFetchError`/`RecipeSourceParseError` thrown by
|
|
* the generic `jsonLdRecipeAdapter` (`sourceKey` `"jsonLdRecipe"`) as having
|
|
* come from this adapter instead (`sourceKey` `"mangerBouger"`) — same
|
|
* reasoning as `marmiton.ts`/`750g.ts`'s identically-named helpers.
|
|
*/
|
|
function rekeySourceError(err: unknown): unknown {
|
|
if (err instanceof RecipeSourceFetchError) {
|
|
return new RecipeSourceFetchError(SOURCE_KEY, err.message, { cause: err.cause });
|
|
}
|
|
if (err instanceof RecipeSourceParseError) {
|
|
return new RecipeSourceParseError(SOURCE_KEY, err.message, { cause: err.cause });
|
|
}
|
|
return err;
|
|
}
|
|
|
|
/**
|
|
* mangerbouger.fr ("La Fabrique à Menus") — Santé publique France's public
|
|
* nutrition site. Unofficial (`official: false`): no published API, same
|
|
* reasoning as every other adapter in this family — fetching ordinary pages
|
|
* and reading data the site never committed to a stable contract, not a
|
|
* maintained endpoint. `fetchDetail` delegates straight to
|
|
* `jsonLdRecipeAdapter`; `parse` wraps it with {@link patchRecipeJsonLd}
|
|
* (see that function's doc comment for the two real gaps it fixes).
|
|
* `list()` doesn't use JSON-LD at all — see `LIST_URL`'s doc comment.
|
|
*/
|
|
export const mangerBougerAdapter: RecipeSourceAdapter<{ html: string; url: string }> = {
|
|
key: SOURCE_KEY,
|
|
name: "Manger Bouger",
|
|
official: false,
|
|
// Chemin fixe (pas d'icône versionnée/hashée comme sur d'autres sources
|
|
// de cette famille) — répond correctement sans paramètre supplémentaire.
|
|
iconUrl: "https://www.mangerbouger.fr/manger-mieux/la-fabrique-a-menus/favicon.ico",
|
|
// Le contenu de mangerbouger.fr (noms, ingrédients, instructions) est en
|
|
// français — détermine contre quel modèle/locale d'étiquettes
|
|
// d'ingrédients translateRecipe (recipe-translation.ts) résout les
|
|
// recettes de cette source lors d'une prévisualisation/d'un import.
|
|
locale: "fr",
|
|
|
|
async list(params: RecipeSourceListParams): Promise<RecipeSourceListResult> {
|
|
try {
|
|
const page = params.cursor ? Number(params.cursor) : 1;
|
|
const query = params.query ?? "";
|
|
const listUrl = `${LIST_URL}?diet=ALL&page=${page}&query=${encodeURIComponent(query)}`;
|
|
|
|
let response: Response;
|
|
try {
|
|
response = await fetch(listUrl);
|
|
} catch (cause) {
|
|
throw new RecipeSourceFetchError(SOURCE_KEY, `Network error listing recipes (${listUrl})`, {
|
|
cause,
|
|
});
|
|
}
|
|
if (!response.ok) {
|
|
throw new RecipeSourceFetchError(
|
|
SOURCE_KEY,
|
|
`mangerbouger.fr responded ${response.status} (${listUrl})`,
|
|
);
|
|
}
|
|
const html = await response.text();
|
|
|
|
const data = extractNextData<MangerBougerListPageData>(html);
|
|
const state = data?.props?.initialState?.recipes;
|
|
|
|
const items: RecipeSourceListItem[] = (state?.list ?? [])
|
|
.filter((entry): entry is MangerBougerListEntry & { slug: string; name: string } =>
|
|
Boolean(entry.slug && entry.name),
|
|
)
|
|
.map((entry) => ({
|
|
externalId: detailUrl(entry.slug),
|
|
title: entry.name,
|
|
picture: entry.image ?? null,
|
|
url: detailUrl(entry.slug),
|
|
}));
|
|
|
|
return { items, nextCursor: state?.hasMorePages ? String(page + 1) : null };
|
|
} catch (err) {
|
|
// Rethrown as-is (already keyed "mangerBouger" by whichever branch
|
|
// above threw it) — this adapter's only caller (`sources.service.ts`)
|
|
// already handles/logs failures centrally; this method just isn't
|
|
// allowed a bare `async` body without a try/catch per the repo's
|
|
// convention. Same reasoning as `marmiton.ts`/`750g.ts`.
|
|
throw err;
|
|
}
|
|
},
|
|
|
|
// `externalId` est directement l'URL canonique de la recette sur
|
|
// mangerbouger.fr (renvoyée telle quelle par `list()` ci-dessus) — même
|
|
// convention que `jsonLdRecipeAdapter.fetchDetail`, à qui cette méthode
|
|
// délègue entièrement : la réparation du JSON-LD (voir
|
|
// `patchRecipeJsonLd`) n'a lieu qu'à l'étape `parse()`, pas ici.
|
|
async fetchDetail(externalId: string): Promise<{ html: string; url: string }> {
|
|
try {
|
|
return await jsonLdRecipeAdapter.fetchDetail(externalId);
|
|
} catch (err) {
|
|
// Pas un simple re-throw : `rekeySourceError` est le traitement utile
|
|
// que ce point d'appel doit faire de l'erreur (relabelliser sa
|
|
// `sourceKey`), conformément à la convention await/try-catch du repo.
|
|
throw rekeySourceError(err);
|
|
}
|
|
},
|
|
|
|
parse(raw: { html: string; url: string }): ParsedRecipe {
|
|
try {
|
|
const patchedHtml = patchRecipeJsonLd(raw.html);
|
|
return jsonLdRecipeAdapter.parse({ html: patchedHtml, url: raw.url });
|
|
} catch (err) {
|
|
throw rekeySourceError(err);
|
|
}
|
|
},
|
|
};
|