#95 — Vorratsschrank — KI-Erkennung unbekannter Produkte per Foto/Sprache (Phase 2) #95
Labels
No labels
priority/could
priority/must
priority/should
priority/wont
status/blocked
status/claimed
status/done-migrated
type/bug
type/feature
type/infra
type/tech-debt
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
robert/todo#95
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Story: Vorratsschrank — KI-Erkennung unbekannter Produkte per Foto/Sprache (Phase 2)
As a Haushaltsmitglied,
I want to ein Produkt, dessen Barcode in keiner Datenbank bekannt ist, per Foto oder Spracheingabe statt
reiner Texteingabe identifizieren lassen,
so that ich unbekannte Produkte nicht jedes Mal komplett von Hand eintippen muss.
Depends on:
#94(Phase 1 — Barcode-Scan, manuelle Text-Eingabe als Fallback existiert bereits).Acceptance criteria:
Wissensdatenbank bekannt, bietet die App zusätzlich zur bestehenden Text-Eingabe zwei weitere Eingabewege
an: Foto und Spracheingabe.
Produktnamen-Vorschlag liefert; der Nutzer bestätigt oder korrigiert ihn.
ebenfalls bestätigt/korrigierbar.
zweimal gefragt).
Out of scope for this story:
dem Menschen (Kosten, Datenschutz).
#82(rein lokal im Browser) ist eine externe Anbindung hier bewusst akzeptiert,da lokale Modelle für generische Foto-/Spracherkennung in dieser Qualität im Browser aktuell nicht realistisch
sind.
Sicherheits-Vorprüfung ist für diese Story verpflichtend, nicht optional: Dies wäre die erste Funktion der
App, die Nutzerdaten (Fotos aus der eigenen Küche, Sprachaufnahmen) an einen externen Drittanbieter sendet.
Bevor implementiert wird, muss die Security-Agent-Vorprüfung (siehe
ai/roles/00_team_overview.md,Feature-Zyklus) explizit klären: welcher Anbieter, welche Daten die App verlassen, ob sie beim Anbieter
gespeichert/für Training verwendet werden, und ob vor der ersten Nutzung ein Zustimmungs-/Datenschutzhinweis
nötig ist.
Open questions:
Gibt es eine bevorzugte KI-Anbindung (z. B. bereits vorhandener Account/API-Key), oder soll der ArchitectResolved by human, 2026-08-11:frei vorschlagen? → Rückfrage an den Menschen bei Ausarbeitung.
"Please use an common standard, and make the url and api key configureable. It should work with local
ollama as well." → OpenAI-compatible HTTP API (
/chat/completionsfor vision,/audio/transcriptionsfor speech), base URL + API key + model names configurable via deployment config (env vars/
appsettings.json, same pattern as SMTP/push VAPID), not a per-user setting. See95_pantry_ai_product_recognition_phase2_design.mdfor what this does and doesn't cover (Ollama servesthe vision half natively; its own text-generation server has no audio-transcription endpoint, so speech
needs a different self-hosted or cloud endpoint that speaks the same Whisper-compatible API shape).
design (
95_pantry_ai_product_recognition_phase2_design.md)Design:
#95— Vorratsschrank KI-Erkennung unbekannter Produkte per Foto/Sprache (Phase 2)Architect note (2026-08-11). Human resolved the story's own open question: use a common,
self-hostable HTTP standard rather than a single hardcoded vendor SDK, with the endpoint, key, and model
names all deployment-configurable, and Ollama specifically supported. That points at the OpenAI-compatible
API shape — the de facto standard every major self-hosted inference server (Ollama, LocalAI,
llama.cpp's server, vLLM) and every cloud vendor with an "OpenAI-compatible" mode already implements:
POST {base}/chat/completionsfor vision (image content parts),POST {base}/audio/transcriptionsforspeech-to-text (Whisper-shaped multipart upload). One
HttpClient, one base URL, one API key, twoindependently-configurable model names.
Scope carried over from Phase 1
#94'sScanPantryProductBarcodeCommandalready has the exact mechanism the story's 4th AC asks for("Das bestätigte Ergebnis wird wie in Phase 1 dauerhaft mit dem Barcode verknüpft gespeichert, nie
zweimal gefragt"): when a barcode is unknown to both the pantry's own products (dedup by
(PantryId, Barcode)) and Open Food Facts, the caller suppliesFallbackNameand a newPantryProductEntitypermanently linked to that barcode is created. Phase 2 does not touch thatpersistence path at all — it only adds two new ways to produce the string the user was already typing
by hand: a photo-recognition suggestion and a speech-transcription suggestion, both still shown to the
user for confirmation/correction before the existing submit path runs. Same for Shopping's
ScanShoppingProductBarcodeCommand(#115already shares Open Food Facts lookup with Pantry; this extendsthat sharing to the new AI lookups too).
New deployment config —
AiSettingsTwo independent
boolflags, not one:IsPhotoRecognitionConfigured/IsSpeechRecognitionConfigured,each requiring
BaseUrland its own model name. This is a deliberate design response to the human'sown "should work with local Ollama" requirement — Ollama's OpenAI-compatible layer serves
/chat/completions(including vision models) but has no/audio/transcriptionsendpoint at all. Anoperator running Ollama-only sets
VisionModeland leavesTranscriptionModelblank; they get the photooption only, not a "Sprache" button that always 501s. Same env/appsettings/docker-compose wiring pattern
as
Email/Push(.env.exampledocumentsAI_BASE_URL/AI_API_KEY/AI_VISION_MODEL/AI_TRANSCRIPTION_MODEL, mapped throughdocker-compose.yml'sAi__BaseUrletc.). NoValidateOnStart()— unlikeAppSettings.FrontendBaseUrl, "all blank" is a valid, fully-supported"feature off" state, not a misconfiguration.
New backend surface
IAiProductRecognitionClient(CqsTodo/Ai/), oneHttpClient-backed implementation:TryRecognizeProductFromPhoto(byte[] jpegBytes, CancellationToken)— one-shot chat completion, afixed system prompt constraining the model to respond with a short grocery-product name and nothing
else (no free-form chat, no injected user text — the only user-controlled input is the image itself),
image sent as a
data:image/jpeg;base64,...content part per the OpenAI vision message shape.TryTranscribeSpeech(byte[] audioBytes, string contentType, CancellationToken)— multipart upload to/audio/transcriptions(file,model), returns the transcript text verbatim (the frontend dialogstill shows it in an editable field — a mis-transcription is just a wrong prefill, not a data-integrity
issue, same trust level Open Food Facts' suggestion already has).
throw out to the caller — any HTTP/timeout/malformed-response failure is logged and swallowed to
null, matchingOpenFoodFactsClient's own "external outage never breaks the feature" convention.Whichever modality's suggestion comes back
null, the user still has the plain manual-entry text fieldas the ultimate fallback — nothing about this feature can make entering a product name harder than
Phase 1 already made it.
under either feature folder):
GetAiRecognitionCapabilitiesQuery() -> AiRecognitionCapabilitiesDto(bool PhotoSupported, bool SpeechSupported)— lets the frontend hide a button that would always fail instead of discovering thatby calling it and getting
nullback.RecognizeProductNameFromPhotoQuery(string ImageBase64) -> string?RecognizeProductNameFromSpeechQuery(string AudioBase64, string ContentType) -> string?WithAuthorization(_ => new AuthorizeIsCurrentUserAuthenticatedQuery())— login-gated likepush subscriptions/API keys, but deliberately not pantry/shopping-list-scoped, since recognizing
"what's in this photo" needs no access to a specific list's data; the existing
AuthorizePantryAccessForCurrentUserQuery/AuthorizeShoppingListAccessForCurrentUserQuerydecoratorsstill gate the actual
Scan*Commandthat persists the confirmed name against a barcode.UpdateUserAvatarCommandHandler.ProcessUpload's decode/validate core (base64 parse, size cap, formatallowlist via
Image.Identify, dimension cap, animated-frame rejection,UnknownImageFormatException/InvalidImageContentExceptionhandling) into a sharedImageUploadValidator.LoadAndValidate(base64, maxBytes, maxDimension)returning a validated, decodedImage; each caller does its ownencode/resize afterwards (avatar: crop-to-square 256px WebP; AI photo: cap-to-1024px-longest-side JPEG,
no cropping — cropping a kitchen product photo could cut off the exact label text the model needs to
read). Same two guards apply to the AI path as already apply to avatars: a decompression-bomb-shaped
file is rejected before the full decode by dimension-checking by
Image.Identifyfirst, and there-encode step means whatever EXIF/metadata (including GPS, if a phone attaches it) the original photo
carried is never forwarded to the external AI endpoint — only re-encoded pixel data is.
AudioUploadValidator) — no existingdecode/re-encode path to extend, and no audio-processing library in this repo's dependency set, so this
intentionally does not re-encode audio (there's nothing here to strip —
MediaRecorder-producedWebM/Ogg has no EXIF-equivalent metadata risk). Enforces a size cap (10 MB — a few minutes of
browser-recorded Opus easily fits, avoids a large-request DoS surface, similar reasoning to the photo's
2 MB cap) and a magic-byte content-type sniff against an allowlist (WebM/Ogg/WAV/MP3 signatures) so a
client-forged
Content-Typeclaim can't smuggle an arbitrary file past the size check into the outboundmultipart request to the external endpoint.
New frontend surface
usePhotoCapturehook (ReactUi/src/hooks/), a sibling touseBarcodeScannerrather than anextension of it — found while implementing that by the time a user reaches the "unknown product"
fallback,
useBarcodeScanner's own successful-decode callback has already calledcontrols.stop(),so there's no still-open stream left to snapshot from (the design's original plan to extend that hook
didn't hold up against its actual lifecycle).
usePhotoCaptureopens its owngetUserMediavideostream (no second permission prompt — browsers grant camera access per origin, not per call) and
exposes
capturePhoto(): string | null, which draws the current frame to an offscreen<canvas>andreturns a base64 JPEG.
useVoiceRecorderhook (ReactUi/src/hooks/):getUserMedia({ audio: true })+MediaRecorder, exposing{ isRecording, start, stop, error };stop()resolves the recorded Blob asbase64 via
FileReader. FirstMediaRecorder/audio-getUserMediausage in this codebase (confirmedvia repo search — no prior audio capture existed), sibling to the existing camera hook rather than a
generalized "media capture" abstraction, since the two have different lifecycles (continuous decode loop
vs. start/stop-once recording).
UnknownProductNameEntrycomponent (ReactUi/src/components/) replaces the near-identicalmanual-
<Input>-only fallback block duplicated today inPantryBarcodeScanner.tsxandShoppingBarcodeScanDialog.tsx: the text input (unchanged), plus a " Foto" button (shown only whenGetAiRecognitionCapabilitiesQuery().PhotoSupported) and a " Sprache" button (shown only when.SpeechSupported), each producing a suggestion that fills the same editable input the user alreadyconfirms/corrects before submitting — worth extracting now specifically because Phase 2 adds real
behavior (camera capture, recording, two new network calls, a privacy notice) to what was one
<Input>;copy-pasting that into both call sites would have tripled the new logic instead of sharing it.
toast — see security pre-review point 3 below for why): "Foto/Aufnahme wird an den konfigurierten
KI-Dienst gesendet."
Out of scope (per story)
No admin/settings UI for the AI config itself — this repo has no precedent anywhere for exposing
deployment-level external-service config in the frontend (SMTP/push VAPID are both env-only, confirmed by
grep), and the human's own answer asked for env/URL+key configurability, not a UI. No vendor selection UI.
No retry/fallback chain across multiple configured providers — one configured endpoint, full stop; an
operator who wants a specific provider's exact behavior configures that provider's URL directly.
security_prereview (
95_pantry_ai_product_recognition_phase2_security_prereview.md)Security Pre-Review:
#95— Vorratsschrank KI-Erkennung unbekannter Produkte per Foto/Sprache (Phase 2)Reviewed before implementation, per
ai/roles/00_team_overview.md's feature cycle (design → securitypre-review → implementation) and this story's own explicit "Sicherheits-Vorprüfung ist für diese Story
verpflichtend, nicht optional" clause. This is the first feature in the app that sends user-originated
photos/audio to an external third party, so the story's own required questions are answered directly
below before anything else.
Story's own required questions
Welcher Anbieter? None fixed — the human's resolution deliberately made this a deployment choice, not
a build-time one (
Ai:BaseUrl/Ai:VisionModel/Ai:TranscriptionModel, blank = feature off). Whoever runsthis app instance chooses their own endpoint: a self-hosted Ollama (data never leaves their own
infrastructure), a self-hosted Whisper-compatible transcription server, or a cloud vendor with an
OpenAI-compatible API.
Welche Daten verlassen die App? A JPEG-re-encoded photo (stripped of any original metadata/EXIF/GPS —
see point 2) when "Foto" is used, or a WebM/Ogg/WAV/MP3 audio clip when "Sprache" is used. Nothing else —
no session cookie, no user id, no pantry/list identifiers are included in either outbound request (see
point 4).
Werden sie beim Anbieter gespeichert/für Training verwendet? Unknowable by this app at build time
— it depends entirely on whichever endpoint the operator configures, and is exactly the kind of fact this
app cannot verify or enforce technically. This is answered by disclosure, not by code: the in-app notice
(point 3) tells the user that their photo/recording leaves the app to "the configured AI service" — it
is the deploying operator's responsibility (same as choosing an SMTP relay or a Seq log sink today) to
pick an endpoint whose data-handling policy they're comfortable with, and to document that for their own
users if it differs from "processed only, never retained." This app does not claim a specific vendor's
policy, since it doesn't know which vendor will be configured.
Braucht es vor der ersten Nutzung einen Zustimmungs-/Datenschutzhinweis? Yes — required, not
optional. See point 3 for why it must be always-visible rather than a one-time dismissible dialog.
1. Blast radius of a "wrong" AI response
Risk: A malicious or malfunctioning configured endpoint returns attacker-controlled text as the
"recognized" product name or transcript.
Mitigation: The returned string is never trusted further than Phase 1 already trusts Open Food
Facts' response (see
94_..._security_prereview.mdpoint 5) or a user's own typed input: it only everprefills the same editable text field the user must still confirm before submitting, and on submission it
passes through the exact same
PantryProductName/ShoppingProductNameVogen validator (length cap, nospecial-casing) as manual entry. It is rendered only as escaped React text, never as markup/HTML. It can
never reach a SQL query, a shell command, or any privileged action — the only thing "downstream" of this
string is a plain varchar column.
2. Photo upload — decompression bombs, format spoofing, metadata leakage
Risk: Same class of risk as the existing avatar upload (
UpdateUserAvatarCommandHandler): anoversized or maliciously-crafted image could exhaust server memory during decode, or forward
identifying metadata (GPS EXIF from a phone photo of the kitchen) to the external AI endpoint.
Mitigation: Reuses the avatar path's exact validation core (extracted into
ImageUploadValidator.LoadAndValidate): base64-decode with a byte-length cap before any image decode,Image.Identify(header-only, cheap) checked against dimensions and an allowlist of decodable formatsbefore the full
Image.Load, and animated multi-frame images rejected — all before the expensive fulldecode ever runs, closing the same "many-large-frames-in-a-small-file" bomb vector the avatar review
already closed. The photo is then always re-encoded to plain JPEG before being sent onward — the original
file's bytes (and anything embedded in them) are discarded entirely; only decoded pixel data survives the
round-trip, so no EXIF/GPS/embedded-metadata of any kind reaches the external endpoint even if the source
photo carried it.
3. Audio upload — format spoofing, size DoS
Risk: An arbitrarily large or mislabeled file forwarded to an external endpoint as "audio" (resource
exhaustion, or smuggling an unrelated file type past a naive
Content-Type-trusting check).Mitigation: Size cap enforced on the decoded byte length before any further processing (10 MB — well
above a realistic few-minutes-long browser voice recording). Content type is verified by sniffing the
first bytes against known container signatures (WebM/EBML, OggS, RIFF/WAVE, MP3 frame sync/ID3) rather
than trusting the client-supplied MIME string, so a forged
Content-Typeheader can't bypass theallowlist. No re-encoding is done (no audio library in this repo's dependencies, and browser-recorded
WebM/Opus carries no EXIF-equivalent identifying metadata the way a phone photo does) — the size cap and
signature check are the full mitigation here, proportionate to the actual risk shape.
4. Outbound request — SSRF surface, credential handling
Risk: Unlike Open Food Facts' fixed, code-constant base URL,
Ai:BaseUrlis operator-configurable —but a config value the operator sets themselves (like
Email:Smtp:HostorPush:Vapid:*already are) isa fundamentally different trust boundary than user-supplied input. No part of the base URL/host is ever
derived from a request the frontend sends — a logged-in user cannot redirect the outbound call anywhere,
regardless of what they submit as
ImageBase64/AudioBase64. This is the same trust model this codebasealready applies to the DB connection string and SMTP host: an operator with config-file/environment access
is not an untrusted party in this app's threat model.
Mitigation:
Ai:ApiKey, when set, is sent only as anAuthorization: Bearerheader to theconfigured
Ai:BaseUrl— never logged (handlers log request/response shapes — "recognitionsucceeded/failed" — never the key, the image bytes, the audio bytes, or the returned text verbatim beyond
what's already user-visible in the UI it prefills). No session cookie, auth token, user id, or any other
app-internal identifier is included in the outbound request to the AI endpoint — the request carries only
the media bytes and a fixed system prompt, so a compromised or malicious configured endpoint learns
nothing about the app's users beyond the content of whatever they chose to photograph/say. 30s timeout +
try/catch around the whole call, same "external outage can never break or hang the feature" convention as
OpenFoodFactsClient.5. Prompt injection via image/audio content
Risk: A crafted image (text overlay) or spoken phrase could attempt to make the underlying model
ignore its system prompt and return something other than a plausible product name — e.g. an
attacker-controlled long string.
Mitigation: Bounded blast radius, not prevented outright (prompt injection against a
third-party-hosted model isn't something this app's own code can fully close, and doesn't need to be —
see point 1): the model's entire "authority" is a single string returned into a user-editable text field
that the user reviews before it becomes a
PantryProductName/ShoppingProductName, capped by that valueobject's existing length validation exactly like every other name in this app, manual or suggested. The
worst case is a nonsense or overlong-and-truncated suggestion the user simply retypes — functionally
identical to Open Food Facts returning a bad name today, already an accepted, non-blocking risk shape in
this codebase (
94_..._security_prereview.mdpoint 5).6. Consent notice — why always-visible, not a one-time dialog
Risk: A dismiss-once "I understand" dialog (like a typical cookie banner) would stop being accurate
information the moment an operator changes their configured endpoint — a user who dismissed it once
under "data stays local (Ollama)" wouldn't be re-notified if the operator later switches to a cloud vendor.
Mitigation: The notice is a fixed, always-rendered line directly above the Foto/Sprache buttons in
UnknownProductNameEntry, not a dismissible interstitial — it costs nothing to show every time (this flowis already a rare, occasional "unknown barcode" path, not a hot path where repetition would be
disruptive), and never goes stale relative to whatever endpoint happens to be configured right now.
Ergebnis
No blocking findings. Point 4's SSRF concern is resolved by trust-boundary reasoning (operator config,
not user input) rather than technical prevention, consistent with how this codebase already treats every
other operator-supplied external-service endpoint. Approved for implementation as designed.
security_final (
95_pantry_ai_product_recognition_phase2_security_final.md)Security Final Review:
#95— Vorratsschrank KI-Erkennung unbekannter Produkte per Foto/Sprache (Phase 2)Performed after implementation, per
ai/roles/00_team_overview.md's feature cycle. A dedicatedsecurity-review subagent traced authorization wiring, upload validation, the outbound HTTP call, and
data flow for the returned AI suggestion end to end. No high-confidence findings.
Findings
None. Verified specifically:
GetAiRecognitionCapabilitiesQueryHandler,RecognizeProductNameFromPhotoQueryHandler,RecognizeProductNameFromSpeechQueryHandler) wireWithAuthorization(_ => new AuthorizeIsCurrentUserAuthenticatedQuery()), identical to theestablished bare-login-gate pattern (
GetPushVapidPublicKeyQueryHandlerand 30+ siblings). Nonetouch per-user/per-list data, so there is no cross-user data exposure.
ImageUploadValidator.LoadAndValidatere-decodes bytes independent of anyclaimed content type, enforces the size cap, checks the actual decoded format against an allowlist,
bounds pixel dimensions, and rejects animated images before the full decode.
AudioUploadValidatordoes real magic-byte container sniffing (WebM/Ogg/WAV/MP3), never trusting a client-supplied MIME
string.
UpdateUserAvatarCommandHandler.csbefore/after theImageUploadValidatorextraction — the exact same size cap (2 MB), format allowlist,MaxSourcePixelDimension(4096), and animated-frame rejection survived verbatim. No regression fromthe refactor.
Ai:BaseUrlis read only from server configuration at DI-registrationtime; no user-supplied value ever reaches
HttpClient.BaseAddressor influences protocol/host.concatenated into it — only image bytes are attacker-influenced. The AI's returned suggestion flows
into a React-controlled
<Input>(nodangerouslySetInnerHTMLanywhere in the changed files) andstill passes through the same Vogen validators as manually-typed input before persistence.
AiProductRecognitionClientlogs only exception objects on failure,never request/response bodies. The AI API key is only ever added as an
Authorization: Bearerheaderon the dedicated outbound
HttpClient, never logged or returned to the frontend —AiRecognitionCapabilitiesDtoexposes only two booleans, not the base URL or key.Result
No blocking findings. Approved as implemented. See
95_pantry_ai_product_recognition_phase2_security_prereview.mdfor the pre-implementation review thisfollows up on.