A working research prototype — TRL 3, running live in this tab.
It answers from authorities it has been loaded with and cites them, and
declines the rest rather than guessing. It is not a clinical tool, not
legal advice, and not a case-management system.
The engine downloads once — about 22 MB,
near 19 MB compressed over the wire — and your browser then caches it.
While it answers, it also fetches the law it reasons over in the
background: six citation and provenance vocabularies (under 1 MB
together) and Title 42 of the U.S. Code, about 35 MB. Ask
throughout — more becomes answerable as each arrives, the source panel
shows what has landed, and you can put any of it back down. After that it
answers with no network at all, and nothing you type leaves this device.
A working research prototype — TRL 3, running live in this tab.
It answers questions about what U.S. federal caregiving and HCBS
rules say, from loaded authorities it cites, and declines
the rest. It is not a clinical tool, not legal advice, and not a
case-management system. Everything below is computed here, now —
including the count of what it does not yet answer.
Every answer arrives with the provision it came from
A reasoning engine that answers family-caregiver and HCBS/EVV workforce-compliance questions from loaded statute and regulation, and shows you the governing provision on every one, so you can open it and check. Where no loaded authority governs a question, it names what it could not ground, so the question reaches a person already sharpened. Everything below runs live, in this browser, against the real evaluation corpus.
What you can check in the next sixty seconds
It cannot compose a citation it was never given: there is no generative stage, so an answer exists only where a loaded, cited provision does. Hand it an invented statute and watch it decline.
Deterministic — no sampling, no stochastic inference. Ask the same question twice and the two answers are identical, today and next quarter.
The whole engine runs in this browser tab via WebAssembly. No server, no account, no GPU, and no question you type ever leaves this device.
Every number on this page is computed here, now, or shows the exact test command that re-derives it.
Starting worker…
The engine downloads once — about 22 MB, near 19 MB over the wire compressed — and is then cached by your browser. Once it is answering, it also fetches the statutory text it reasons over in the background: six citation and provenance vocabularies (under 1 MB together) and Title 42 of the U.S. Code, about 35 MB, which carries Medicare, Medicaid, the Older Americans Act and the HCBS waiver authorities. You can ask questions throughout; more becomes answerable as each arrives, and the source panel shows exactly what has landed. After that it answers with no network at all, and no question you type leaves this device.
—
concepts loaded
—
ontologies loaded
—
Bank real questions collected so far — the bank grows as more are collected
—
Done answered today from a governing provision, citation attached — a floor a committed test holds
This track's slice of the evaluation instrument, and what the engine grounds in it right now. Copy any question and paste it into the Chat tab to see it answered live, with the provision attached.
✅ Grounded questions in this track
Drawn at random from the Track 1 rows the engine answers today. Each returns a definition the engine already holds, with what grounded it shown; the two purpose-built lexicons carry the authority that defines the term inline.
Track 2 — pr4xis HCBS/EVV Compliance Navigator
What a compliance staffer can ask today
This track's slice of the evaluation instrument, and what the engine grounds in it right now. Copy any question and paste it into the Chat tab to see it answered live, cited to the federal rule.
✅ Grounded questions in this track
Drawn at random from the Track 2 rows the engine answers today. Each returns a definition or rule carried from the provision that governs it.
Questions to try
Copy any question below, open the Chat tab, and paste it. The same engine answers it live in this browser.
✅ Questions it answers today
Drawn uniformly at random from the rows the engine grounds right now. Each one returns an answer the engine already holds, showing what grounded it; definitions from the two purpose-built lexicons carry the provision they were authored from, so you can open it and check. Draw again for a different three.
🛑 Questions built to trick it
Invented statutes, fabricated program names, false premises. There is no generative stage, so there is nothing to compose an authority out of — copy any of these into Chat and it comes back declined. Draw again for a different three.
Engine Evidence Lab
Transparency into how the engine performs across the real evaluation corpus — Track 1, Track 2, or combined. Every number below is computed live from the corpus fetched by this page load; nothing is hardcoded.
corpus: —
Honesty & safety
Each check runs live, in this browser, against the real engine — click "Verify live" to re-derive it yourself.
What we are building next — each one named, tracked, and under a committed CI bound
Every number below is a real, named, reproducible list of specific questions — not an estimate. Closing these is live, ongoing work (see the repository's own commit history for the loop), not a static disclaimer.
A few real examples, drawn once per visit to this tab from the committed corpus snapshot:
Sample pipeline trace
The real typed trace from the last question you ran in the Chat tab.
Ask a question in the Chat tab to populate this.
Smart-40 Validation Protocol (as submitted to ACL)
Loading the published protocol…
Method, evidence, and the judging map
How the ACL rubric's six categories, and the seven AI Principles within category five, map onto real evidence in this system — for a time-pressured reviewer.
The reasoning pipeline
Every turn runs through these ontologies in order. They are compiled into the engine rather than loaded at runtime, which is why they do not appear in the source catalog — that lists the vocabularies you can load and unload. Read here from the same constants the per-answer trace reports.
Model card
What it is
A neurosymbolic reasoning engine. Knowledge is typed ontologies connected by category-theory functors. Every turn ends in one of four typed outcomes: an answer carrying the authority it was built from, a cited rule together with the one fact it still needs, that rule resolved once the fact arrives, or a decline. A term outside the two purpose-built lexicons answers from a general-vocabulary sense, labelled as such and tracked as a named defect in the Track 1 appendix.
Intended use
Answering caregiving and HCBS/EVV compliance questions from loaded, cited definitions. Not medical, legal, or benefits advice.
Out of scope
No scheduling, recruitment, training automation, or EMR/EHR data exchange.
Determinism
No sampling, no stochastic inference — identical input yields identical output. Every metric on this page is reproducible from a named test.
Failure mode
When no cited definition is loaded, it declines rather than fabricating an answer, and names the surface it could not ground wherever an unresolved surface exists to name.
AI facts label
Runs on
Your device — WebAssembly, no server, no account.
Data sent off-device
None. Questions never leave the browser.
Training data
Not a trained model. Loads authoritative statute/regulation text and cited definition lexicons.
Sources
USC/CFR, CMS guidance, state EVV FAQs, WordNet — each answer names the ontologies it reasoned over.
Confidence display
Categorical outcomes only (answered / abstained / conditional). No misleading confidence percentages.
Human oversight
Conditional rules ask for the missing fact instead of guessing, and every answer in the Chat tab carries a “Download this decision record (JSON)” link — the typed outcome, the governing rule and its citation, and the full reasoning trace, written in your browser as a file your compliance function keeps.
This page's palette, checked at view time by the engine's own cited colour ontology — computed from the page's actual rendered tokens rather than asserted. This is a contrast check over the palette, not a full Section 508 conformance review.