Musnad Middle Ground — Phase 1 Recap

Musnad Middle Ground: The Goal

Middle Ground’s blog, khuṭbahs, classes, and podcasts quote the same ḥadīth across dozens of pieces of content — but with varying wording, citation quality, translation, and narrator spelling. The Musnad Middle Ground project exists to create one stable, canonical record for each verified report while preserving how it originally appeared.

The core principle that governs everything below:

machine does repetition
human does judgment
deterministic code does mutation
Git preserves accountability

Phase 1 asked a specific question: can an AI agent responsibly assist with finding and cross-referencing ḥadīth citations across the Middle Ground website, without ever crossing into scholarly judgment?


What We Built

1. A Website Audit Skill (musnad-website-audit)

A three-stage pipeline for scanning muslim.center blog posts:

  • Stage 1 — Extract: scan a post for ḥadīth-like passages using both English triggers (“the Prophet ﷺ said,” “it is narrated,” collection names like Bukhārī/Muslim/Abū Dāwūd) and Arabic triggers (حديث, قال رسول الله). Deliberately excludes passages that only discuss the term “ḥadīth” without quoting a narration, and separately tracks citations the post already provides for itself.

  • Stage 2 — Verify: for each candidate, attempt to find the ḥadīth in two major online repositories:

    • sunnah.com — the Nine Books plus other primary collections, searchable by English or Arabic phrase, wildcard, and fuzzy matching
    • dorar.net — a Saudi ḥadīth encyclopedia (الموسوعة الحديثية) that returns not just the text but the grading (ḥasan/ṣaḥīḥ/ḍaʿīf/mawḍūʿ), the narrator, the grading scholar, and — critically — takhrīj: cross-references to every other collection that carries the same narration with its number.

    The search budget is capped at 2–3 attempts per candidate. A miss is recorded as “needs manual research,” not chased further — this is a deliberate scope boundary, not a limitation to fix later.

  • Stage 3 — Report: one markdown file per post, listing every candidate with its possible matches, confidence level (high/medium/low — never “verified”), and any scholarly flags (e.g., a citation graded weak, or a collection attribution that doesn’t match what the post claims).

2. A Live Trial (4 Real Posts)

We ran the skill against four actual muslim.center posts to validate it before trusting it at scale:

Post Candidates found Result
Friday Ghusl, Wuḍūʾ, and the Rewards of Walking to Jumuʿah 2 1 high-confidence match — confirmed the project’s existing canonical mm-h-0001; 1 medium match narrowed to a cluster of related Bukhārī narrations
If You Love God Then Follow Me 3 1 medium match — and caught a real discrepancy: the post attributes a ḥadīth to Muslim, but the matched narration is Bukhārī
What Does Sunnah Really Mean? 1 Found the quoted ḥadīth on dorar.net — but its chain is graded ḍaʿīf (weak) by al-Bayhaqī, with dorar’s own editors flagging an “authentic alternative.” Surfaced for scholarly review, not silently accepted.
Muslims & The Post-Literacy Society 0 new The post already self-cites correctly with a direct dorar.net link — the audit correctly found nothing new to flag

This is the real value of the exercise: the tool didn’t just confirm citations, it caught two things a human proofreader might have missed — a misattributed collection and a weakly-graded chain hiding behind a well-known phrase.

3. A Bot Mode Blueprint (musnad-bot-sop.md)

A five-specialist-bot design for running this work as an ongoing operation inside Hermes’s Bot Mode, rather than one-off manual runs:

  • Coordinator — routes requests, orchestrates the others, escalates anything requiring judgment to a human
  • Extractor — runs the extraction and verification stages (mechanical, no judgment)
  • Scholar — summarizes findings, applies proper Arabic transliteration, prepares evidence packets
  • Validator — processes human decisions against the deterministic Musnad CLI, classifies outcomes as SAFE/BLOCKED/AMBIGUOUS
  • Librarian — manages the shared review workspace, archives completed batches

None of the five bots may decide ḥadīth identity, choose canonical text, verify narrators, or approve publication. Those stay human-only, by design.

4. A Standing Telegram Channel

To make this workflow reachable outside the desktop app, we stood up a dedicated Telegram bot on its own Hermes profile (musnad), separate from the main assistant profile so the two never collide over credentials or shared state.


The Real Obstacle: Local vs. Cloud Models

We initially tried to run the Telegram bot on a local model (Qwen, via LM Studio) to keep costs at zero. This ran into a genuine, unresolved technical snag:

  • LM Studio kept loading the model with a 169,728-token context window — far larger than needed for a chat bot — regardless of explicit --context-length flags passed via the command line, tried multiple ways, across multiple fresh reloads.
  • That oversized context made a single response take roughly five minutes, and separately caused LM Studio to refuse loading a second, larger model at all (“insufficient system resources”), which triggered repeated “Provider unresponsive” errors in the desktop app.
  • We correctly diagnosed why (Auto-Evict and JIT loading were working as designed, but the context-length override wasn’t taking — likely a backend quirk with this particular MLX-format model), but did not solve the underlying LM Studio bug itself.

Decision: rather than keep debugging an open-ended local-inference issue mid-session, we moved the bot to a cloud-first, tiered-fallback configuration:

main:      Anthropic (claude-sonnet-5)
fallback:  OpenAI (gpt-5.5-pro)
fallback:  LM Studio, local (qwen3.8-27b)  — last resort only

This gets fast, reliable responses now, keeps a genuine free/local option as a safety net if both cloud providers ever go down, and leaves the door open to revisit local-only operation once the context-length bug is worth chasing down properly.

Lesson for Phase 2: cost-defrayal through local models is a real and worthwhile goal, but it should be validated before it’s load-bearing for a live tool people depend on — not discovered mid-deployment.


What’s Confirmed Working

  • ✅ The audit skill’s extraction and verification logic, validated against real posts
  • ✅ Live search integration with both sunnah.com and dorar.net (including handling sunnah.com’s Cloudflare protection via browser automation)
  • ✅ A dedicated, isolated Hermes profile for the Musnad Telegram bot
  • ✅ End-to-end message delivery: Telegram → Hermes gateway → cloud model → reply, in seconds
  • ✅ Tiered model fallback configured correctly through Hermes’s GUI (not hand-edited config files)

What’s Not Yet Done

  • The 5-bot Bot Mode team (Coordinator, Extractor, Scholar, Validator, Librarian) exists only as a specification — none of the actual bot profiles have been created yet
  • The website audit has only been run on 4 of roughly 43 posts on the site — the rest are unaudited
  • No scheduled/recurring routine exists yet for the audit (e.g., a weekly automatic re-scan)
  • The LM Studio context-length bug was worked around, not fixed
  • Verified candidates from the audit have not yet been fed into the actual Musnad candidate pipeline (mm-c-* IDs) — by design, the audit produces a report for human review first, not automatic ingestion

Phase 2 Preview

Phase 2 will pick up with:

  1. Creating the five specialist bot profiles per the SOP
  2. Running the website audit across the full site
  3. Wiring a recurring routine (cron) for ongoing audits as new posts are published
  4. Deciding whether audit findings should feed into musnad:prepare-review as a new candidate source, and if so, how

Status: Phase 1 complete. Phase 2 begins next session.