I’m building a small webapp to solve a specific household problem: credit card statements, utility bills, and grocery receipts arrive as PDFs, phone photos, and email attachments, and I want line-item price history over time without manually entering everything into a spreadsheet. This is the technology audit that happened before the first line of code.
The core job is document ingestion with structured extraction: drop a file, extract vendor/date/line items/totals, store it in a way that lets me query “what did carrots cost at Vons last month versus this month” or “how much did I spend on utilities over the past quarter.” Everything else—the dashboard, the charts, the mobile-friendly UI—serves that pipeline.
This post covers the framework choice, the database, the extraction engine, charting, hosting, and the two parts that matter most for financial documents: redaction and retention.
Framework: Next.js, not Ionic
I’ve used Ionic before (the smcg.middleground.dev project), so reusing it was tempting. But Ionic is built for hybrid mobile-first apps that want native device APIs (camera, filesystem, haptics) via Capacitor, with the web as a secondary target. SnapSchaught’s actual requirement is “mobile-friendly web, richer desktop web”—a responsive dashboard, not a native-shell app.
Ionic’s component library is opinionated toward mobile app chrome and fights you when building a “richer, more comprehensive” desktop data view. Multi-panel dashboards, drill-down tables, and dense charts are not what Ionic’s UI kit is optimized for. Reusing it here would be reusing a hammer because I have one, not because this is a nail.
Why not Astro? Astro is exactly right for my blog and the Middle Ground Dev site (static content, near-zero client JS) but wrong here. SnapSchaught is heavily interactive and stateful—live dashboards, drag-drop upload, real-time chart updates, drill-downs—the opposite of Astro’s “ship HTML, hydrate islands sparingly” model.
Why not SvelteKit? A legitimate contender—smaller bundles, less boilerplate, genuinely good DX, and PWA support is solid. I’m not choosing it for one practical reason: the charting and data-grid ecosystem is overwhelmingly React-first (Recharts, TanStack Table, shadcn/ui), and this project leans hard on rich, interactive dashboard components where I want the biggest, best-maintained library selection.
Why Next.js wins:
- App Router gives me file-based API routes for the extraction pipeline (
/api/ingest) alongside the UI in one deploy—no separate backend repo needed - First-class PWA support (
app/manifest.ts+ service worker via Serwist) covers “mobile-friendly, installable, works offline-ish” without needing Capacitor - Native drag-and-drop + File System Access API—no framework-specific gap
- Deploys natively on Vercel (already my house platform)
- The single biggest technical risk is the extraction pipeline, not the frontend framework—pick the framework that gets out of the way fastest and has the deepest library bench for dashboards
Verdict: Next.js App Router + TypeScript + PWA.
Database: Postgres via Neon
The app’s own requirements—item-level history over time, store-vs-store comparison, macro rollups by category/time window—are relational queries with a time dimension. “Average price of X at store Y over time,” “sum spend by category this month vs last”—this is squarely SQL’s job. A real time-series database (TimescaleDB, InfluxDB) is overkill for a single household’s thousands to low-millions of line-item rows.
Why Neon specifically:
- Vercel-native integration—one-click provisioning, env vars auto-wired
- Scale-to-zero pricing—this is a personal app with bursty, low-frequency usage (a few receipts a week, not constant SaaS traffic). Neon’s free tier (100 compute-hours/month, 0.5GB storage) likely covers this indefinitely
- Plain Postgres—more portable and “as open-source as possible” than adopting Supabase’s broader platform (auth, realtime, storage) when I don’t need most of that bundle yet
Schema shape (conceptual):
documents(uploaded file metadata, extraction status, raw text/JSON blob, source type)transactions(one per statement/receipt/bill: date, vendor/store, total, category)line_items(item name, normalized item name, price, quantity, unit, transaction FK, store FK)stores(normalized store names—“Vons #2210” and “Vons” collapse to one entity)categories(groceries, utilities, credit card sub-categories)
The unglamorous but critical design problem: a receipt scan of “ORG CARROTS 2LB” today and “Organic Carrots” next month need to resolve to the same logical item for a price-history chart to mean anything. I’m planning a lightweight normalization pass (fuzzy string matching or LLM-assisted canonicalization) as part of the extraction pipeline, not an afterthought.
Document extraction: LLM-vision, with a self-hosted fallback path
I researched both classic OCR and LLM-vision extraction. Bottom line for this use case:
Classic OCR (Tesseract / PaddleOCR):
- Tesseract: extremely fast, tiny (10MB), runs on a Raspberry Pi, but character-error-prone and has zero semantic understanding—it gives you a bag of text, not “this is a line item with this price.” I’d need hand-built layout/regex heuristics per document type (receipts vs. PDF statements vs. utility bills all format differently).
- PaddleOCR (PP-OCRv5, 2025 release): meaningfully better—near-zero character errors, and its “PaddleOCR-VL” variant is actually a small vision-language model (0.9B params) that outputs structured Markdown/JSON directly. This is the strongest pure open-source option if avoiding LLM API costs/dependency becomes a hard requirement.
LLM-vision extraction (Claude/GPT-4o/Gemini):
- Handles layout semantically out of the box—“this is a receipt, these are line items with prices, this is the store name, this is the date”—with structured JSON output, no per-document-type regex logic needed
- Handles messy real-world receipt photos (crumpled, angled, poor lighting) far more robustly than classical OCR
- Cost and privacy tradeoff: sends financial document images to a third-party API. At household-scale usage (a few dozen documents a month) the cost is trivial, but it’s a real dependency and a real privacy consideration
Recommendation—hybrid, phased:
- Phase 1: Use LLM-vision extraction (Claude via Anthropic API) as the extraction engine. Gets a working, accurate pipeline fastest, validates the schema/UX before investing in self-hosted OCR, and the per-document cost is negligible at personal-use volume.
- Phase 2+: If I want to eliminate the external API dependency for full open-source purity, swap in PaddleOCR-VL (self-hosted, Apache 2.0 licensed) as a drop-in replacement. Because both output the same normalized JSON shape, this is a backend change, not a rearchitecture—build the extraction step as a swappable module from day one.
- Either way, PDFs need a text-layer-first shortcut: many bank/credit-card statement PDFs already contain selectable text (no OCR needed at all)—check for that first with
pdf-parseorpdfjs-distbefore routing to image-based extraction.
Charting: Recharts + TradingView Lightweight Charts
Surveyed the 2026 React charting landscape:
- Recharts—the practical default for React dashboards: composable, well-documented, strong shadcn/ui integration, 3.6M weekly downloads, SVG-based (good for crisp mobile rendering at this app’s data scale)
- TradingView Lightweight Charts—purpose-built for “price over time” line/area charts. Smallest bundle (~12KB gzipped), Canvas-rendered so it stays smooth with frequent updates. Worth pulling in specifically for the “carrots over time” and “store trending” screens where Recharts’ generic line chart would work but Lightweight Charts’ financial-chart polish is a closer fit
Recommendation: Recharts for macro dashboard views (category breakdowns, spend-over-time, pie/donut splits) + Lightweight Charts for micro item-price-history and store-comparison line charts. Both are open-source (MIT), both React-native.
Redaction: before extraction, not after
Redaction needs to be an optional upload-time step: user selects “redact” on a document, the redaction pass runs before the file is persisted to storage and before the original goes to the LLM extraction step.
I already have a Stirling-PDF skill for manual redaction tasks, but it’s too heavy for this VPS (2 CPU / 2GB RAM minimum, running Java + LibreOffice + Tesseract in a persistent container) and architecturally the wrong tool—it’s built for a one-off manual workflow with a human reviewing output, not an automated per-upload pipeline.
Recommended approach: PyMuPDF for PDFs, Presidio Image Redactor for images.
For PDFs—PyMuPDF’s redaction API does this correctly:
page.add_redact_annot(rect)marks a region,page.apply_redactions()deletes the underlying text/image objects from the content stream (not a black box drawn on top—this is the real distinction)- Region targeting can be automatic (regex/pattern search for account numbers, SSN-shaped strings) or manual (user draws a box on a preview)
- Optionally rasterize the redacted page to an image after
apply_redactions()as a belt-and-suspenders step—this guarantees no residual text layer survives - License flag: PyMuPDF is dual-licensed AGPL-3.0 or a paid commercial license. For a personal/household app where I control hosting and I’m not distributing the software to third parties as a closed product, AGPL is not a practical obstacle
For images (JPEG/PNG receipts)—Microsoft Presidio’s presidio-image-redactor: OCR locates text, PII analyzer flags likely-sensitive spans (names, card numbers, SSNs), draws opaque fill rectangles over those boxes before the image is persisted or sent to extraction. MIT-licensed, pure Python + Tesseract.
Pipeline shape:
- User uploads a file and optionally toggles “redact”
- If redaction requested: run the appropriate library call synchronously (sub-second to a few seconds per document)
- Persist only the redacted version if the user wants the sensitive original discarded
- Feed the redacted version to the extraction pipeline—so sensitive data (full account numbers, SSNs) never leaves the boundary at all for documents the user flags as sensitive
Document storage and retention
Where originals live: Vercel serverless has no persistent disk, so object storage is the only real option. Vercel Blob (S3-backed, signed-URL upload flow that lets the browser upload directly without routing through a function, $0.023/GB-month). Keep the DB for metadata + a pointer to the file, not the file itself.
Retention—keep or discard originals after extraction?
Three policies:
- Keep everything indefinitely—largest privacy surface (full statements with account numbers sitting in storage)
- Discard immediately after extraction—best privacy posture, but no way to re-run extraction if a bug is found later, no way to look at the original PDF for a disputed charge
- Recommended: keep only when the user explicitly opts in (a “keep original” toggle at upload, default to discard-after-extraction). Gives me control per-document—keep the mortgage statement I might need for taxes, discard the CVS receipt once its line items are captured
Combine with redaction: if “redact” was selected, store the redacted version as the kept original, never the raw one.
Privacy implications specific to financial documents: Credit card/bank statements contain account numbers, sometimes partial SSNs, full name/address, complete transaction histories. The redaction feature exists specifically to let me strip that before it’s ever persisted or sent to a third-party API. If the LLM-vision extraction path is used on an unredacted document, that document’s image is sent to Anthropic’s API—another reason to make redact-before-extraction the encouraged default.
Hosting plan under middleground.dev
Subdomain: snap.middleground.dev
Vercel fit: follows the exact same pattern as every other property—Next.js/Astro → Vercel → deploy on push to main. New GitHub repo (marcmanley/snapschaught), new Vercel project, added as a subdomain under MGMCUpland (the consolidated team).
Where things run:
- Frontend + API routes: Vercel serverless functions (Next.js API routes)
- Database: Neon, provisioned via Vercel-Neon integration
- File storage: Vercel Blob
- Extraction workers: Vercel serverless function with
maxDurationextended (Pro plan supports up to ~800s) calling the Anthropic API directly—likely fast enough for LLM-vision extraction. Fallback only if self-hosted PaddleOCR is adopted later (needs GPU or sustained CPU, doesn’t fit serverless well—would run as a background worker on the VPS)
Honest framing on “as open-source as possible”
The app’s own code, UI, charting, ORM, and (optionally, phase 2+) OCR can all be fully open-source and self-hostable. The two pieces that aren’t—Vercel Blob and the Claude vision API—are both cleanly swappable (S3-compatible storage; PaddleOCR) if I want to eliminate them later, and both were chosen for phase-1 speed/simplicity, not lock-in.
Worth surfacing this tradeoff explicitly rather than glossing over it.
Next
With the stack chosen, the next step is the phase 1 build plan: repo setup, schema draft, and the task breakdown for actually building this thing.