apyhub
Back
▣ DATA EXTRACTION · DATA VALIDATION

Ingredient Parser From Text API

What it does

The ingredient parser API turns a free-text ingredient list from a food label into structured data. Send text (up to 5,000 characters) with an optional lang hint (it, en, fr, de, es or auto), and get back allergens, additives and dietary signals in a stable JSON format with English labels.

Allergens map to the 14 EU allergen categories, each with an eu_index, a confidence score and a source. Confirmed allergens go in allergens, and "may contain traces of" clauses go in a separate allergens_traces list. Additives come back with their E-number code, official EU name, functional category such as emulsifier, and a recognized flag. The dietary object marks vegan and vegetarian as likely or unlikely and flags alcohol, palm oil and added sugars. Fragments the parser can't match are listed in unrecognized_ingredients. A second endpoint looks up a single E-number such as E322 directly.

Use the ingredient parser API to fill allergen fields in grocery and delivery catalogs, check food labels for compliance before products go live, power allergy and diet filters in shopping apps, and enrich product feeds that carry raw ingredient text.

To get ingredient text off a photographed label, run it through the Image OCR API first. To pull it from a retailer's product page, use the Product Data Extraction API. For nutrition and substitutes per ingredient, add the Ingredient API.

▣ ENDPOINT 01 / 02
POST
Parse an ingredient list
https://api.eu.apyhub.com/ingredients/parse-ingredients/v1/parse

QUICKSTART

GUIDE

Quickstart

Send the ingredient list text and language hint to parse common ingredients and allergens.

curl -X POST "https://api.eu.apyhub.com/ingredients/parse-ingredients/v1/parse" \
  -H "apy-token: $APY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"lang":"auto","text":"Wheat flour, sugar, palm oil, E322 (soy lecithin), milk powder"}'

What you'll get back

Returns a JSON object with these top-level fields: cached (boolean or null), dietary (object), degraded (boolean), additives (array), allergens (array), llm_assisted (boolean), degraded_reason (string or null), allergens_traces (array), and unrecognized_ingredients (array).

Notes: A HINT about the language of the INPUT text, used only to pick which language's dictionary terms to match against. It has NO effect on the response: label values, enum tokens and field names are always in English/stable form regardless of lang. auto (the default) checks all five supported languages and is the right choice whenever the input language isn't reliably known ahead of time. Not implemented, and deliberately out of scope for now: a separate output_lang parameter for localizing label values. The current design (stable id + English label) makes that a pure additive change later — clients that key off id today will not break if output_lang is added.

{
  "allergens": [
    {
      "id": "gluten",
      "label": "Cereals containing gluten",
      "eu_index": 1,
      "confidence": 0.95,
      "source": "dictionary"
    },
    {
      "id": "soy",
      "label": "Soybeans",
      "eu_index": 6,
      "confidence": 0.95,
      "source": "dictionary"
    },
    {
      "id": "milk",
      "label": "Milk",
      "eu_index": 7,
      "confidence": 0.95,
      "source": "dictionary"
    }
  ],
  "allergens_traces": [],
  "additives": [
    {
      "code": "E322",
      "name": "Lecithins",
      "name_it": "Lecitine",
      "category": [
        "emulsifier"
      ],
      "recognized": true,
      "confidence": 1
    }
  ],
  "dietary": {
    "vegan": "unlikely",
    "vegetarian": "likely",
    "contains_alcohol": false,
    "contains_palm_oil": true,
    "contains_added_sugars": true
  },
  "unrecognized_ingredients": [],
  "llm_assisted": false,
  "degraded": false,
  "degraded_reason": null,
  "cached": true
}
TRY ITLIVE · 60 ATOMS
Loading your default key…
The full key is used to call the gateway and stays in this tab — never sent to orbit or saved.
body*
Free-text ingredient list. Maximum 5000 characters — generous for any real product label; longer requests are rejected with 413 before any parsing happens.

About this endpoint

What it does

Parses a free-text ingredient list sent in the request body and returns a structured JSON object with additive, allergen, dietary, and classification fields. The endpoint does not change server state; it analyzes the input text and reports the extracted results.

Request Body

ParameterTypeDescription
langENUMHint about the language of the input text, used only to choose dictionary matching.
Allowed values: auto, it, en, fr, de, es.
Default: auto.
Does not affect response language; labels and tokens remain in English/stable form.
textStringFree-text ingredient list to parse.
Maximum length: 5000 characters.

Response

Returns a JSON object with these top-level fields: cached boolean or null, dietary object, degraded boolean, additives array of objects, allergens array of objects, llm_assisted boolean, degraded_reason string or null, allergens_traces array of objects, and unrecognized_ingredients array of strings.

ParameterTypeDescription
cachedBooleanWhether the response was served from cache. Nullable.
dietaryObjectDietary summary object with stable tokens/booleans.
Fields: vegan (likely / unlikely), vegetarian (likely / unlikely), contains_alcohol (boolean), contains_palm_oil (boolean), contains_added_sugars (boolean).
dietary.veganENUMAllowed values: likely, unlikely.
dietary.vegetarianENUMAllowed values: likely, unlikely.
dietary.contains_alcoholBooleanBoolean flag.
dietary.contains_palm_oilBooleanBoolean flag.
dietary.contains_added_sugarsBooleanBoolean flag.
degradedBooleantrue if the applicative spend cap was reached and the service responded in deterministic-only mode.
additivesObject ArrayParsed additives found in the ingredient list. Each item may include code, name, name_it, category, confidence, and recognized.
additives[].codeStringAdditive code such as an E-number.
additives[].nameStringOfficial EU name in English. Nullable.
additives[].name_itStringOfficial EU name in Italian. Nullable.
additives[].categoryString ArrayFunctional class tokens from the regulatory vocabulary. Nullable. An additive can have more than one category.
additives[].confidenceNumberConfidence score.
additives[].recognizedBooleantrue if the code was found in the regulatory table.
allergensObject ArrayAllergens declared as present, not traces-only matches. Each item may include id, label, source, eu_index, and confidence.
allergens[].idStringStable language-independent identifier.
allergens[].labelStringCanonical EU allergen name in English.
allergens[].sourceENUMAllowed values: dictionary, llm.
allergens[].eu_indexIntegerPosition in Annex II of Reg. (EU) 1169/2011.
allergens[].confidenceNumberConfidence score between 0 and 1.
llm_assistedBooleanWhether the LLM layer was used in the parse.
degraded_reasonENUMAllowed values: lifetime_cap_reached, daily_cap_reached, null.
allergens_tracesObject ArrayAllergens mentioned only in a “may contain traces of” clause. Same item shape as allergens.
allergens_traces[].idStringStable language-independent identifier.
allergens_traces[].labelStringCanonical EU allergen name in English.
allergens_traces[].sourceENUMAllowed values: dictionary, llm.
allergens_traces[].eu_indexIntegerPosition in Annex II of Reg. (EU) 1169/2011.
allergens_traces[].confidenceNumberConfidence score between 0 and 1.
unrecognized_ingredientsString ArrayIngredient fragments that could not be classified, echoed back verbatim from the submitted text.

Notes

allergens contains only confirmed allergens, while allergens_traces is reserved for “may contain traces of” matches. If you want a conservative view, read both arrays; if you only care about confirmed presence, read allergens alone.

▣ ENDPOINT 02 / 02
GET
Direct E-number lookup (zero LLM)
https://api.eu.apyhub.com/ingredients/parse-ingredients/v1/additives/:code

QUICKSTART

GUIDE

Quickstart

Look up an E-number by code, using the code in the URL path.

curl -X GET "https://api.eu.apyhub.com/ingredients/parse-ingredients/v1/additives/:code" \
  -H "apy-token: $APY_TOKEN"

What you'll get back

Returns a JSON object with these top-level fields: code (string), name (string or null), name_it (string or null), category (array of strings or null), confidence (number), and recognized (boolean).

{
  "code": "E322",
  "name": "Lecithins",
  "name_it": "Lecitine",
  "category": ["emulsifier"],
  "confidence": 1,
  "recognized": true
}
TRY ITLIVE · 10 ATOMS
Loading your default key…
The full key is used to call the gateway and stays in this tab — never sent to orbit or saved.

About this endpoint

What it does

Looks up a single E-number by its code path parameter and returns the regulatory data we have for that code. The response includes the code itself, the official EU name, optional Italian name, functional category list, a confidence value, and whether the code was recognized.

Path Parameter(s)

AttributeTypeDescription
codeStringThe E-number to look up, such as an E-prefixed code.

Response

Returns a JSON object with code as a string, name as a nullable string, name_it as a nullable string, category as a nullable array of strings, confidence as a number, and recognized as a boolean. This is the success response for the endpoint.

ParameterTypeDescription
codeStringThe E-number code that was looked up.
nameStringOfficial EU name in English. Nullable.
name_itStringOfficial EU name in Italian. Nullable and provided for traceability only.
categoryString ArrayFunctional class(es) as stable snake_case tokens from the canonical Annex I vocabulary. Nullable.
confidenceNumberConfidence value returned by the service.
recognizedBooleanTrue if code was found in the regulatory table, otherwise False.
▣ COMMON ERRORS

Errors any endpoint can return

400bad_request

Required parameter missing or malformed body.

401unauthorized

API key missing, revoked, or not authorized for this service.

429rate_limited

Your plan's per-second rate exceeded. Retry with exponential backoff.

503upstream_busy

Backend temporarily unavailable. Try again in a few seconds.