Pr4xis

Pr4xis

Axiomatic intelligence — runs entirely in your browser Docs
Loading…
ACL Caregiver AI Challenge — Track 1 & Track 2 → Pick up a real family-caregiver or HCBS/EVV compliance question from either track's working set and run it here in Chat, or open the Engine Evidence Lab to watch the published evidence reproduce live in your own browser.
A working research prototype — TRL 3, running live in this tab. It answers from authorities it has been loaded with and cites them, and declines the rest rather than guessing. It is not a clinical tool, not legal advice, and not a case-management system. The engine downloads once — about 22 MB, near 19 MB compressed over the wire — and your browser then caches it. While it answers, it also fetches the law it reasons over in the background: six citation and provenance vocabularies (under 1 MB together) and Title 42 of the U.S. Code, about 35 MB. Ask throughout — more becomes answerable as each arrives, the source panel shows what has landed, and you can put any of it back down. After that it answers with no network at all, and nothing you type leaves this device.
Loading…

Overview

pr4xis Caregiver Answer Engine & HCBS/EVV Compliance Navigator
A working research prototype — TRL 3, running live in this tab. It answers questions about what U.S. federal caregiving and HCBS rules say, from loaded authorities it cites, and declines the rest. It is not a clinical tool, not legal advice, and not a case-management system. Everything below is computed here, now — including the count of what it does not yet answer.
pr4xis · ACL Caregiver AI Challenge, Track 1 & Track 2

Every answer arrives with the provision it came from

A reasoning engine that answers family-caregiver and HCBS/EVV workforce-compliance questions from loaded statute and regulation, and shows you the governing provision on every one, so you can open it and check. Where no loaded authority governs a question, it names what it could not ground, so the question reaches a person already sharpened. Everything below runs live, in this browser, against the real evaluation corpus.

What you can check in the next sixty seconds
  • It cannot compose a citation it was never given: there is no generative stage, so an answer exists only where a loaded, cited provision does. Hand it an invented statute and watch it decline.
  • Deterministic — no sampling, no stochastic inference. Ask the same question twice and the two answers are identical, today and next quarter.
  • The whole engine runs in this browser tab via WebAssembly. No server, no account, no GPU, and no question you type ever leaves this device.
  • Every number on this page is computed here, now, or shows the exact test command that re-derives it.
Starting worker…
The engine downloads once — about 22 MB, near 19 MB over the wire compressed — and is then cached by your browser. Once it is answering, it also fetches the statutory text it reasons over in the background: six citation and provenance vocabularies (under 1 MB together) and Title 42 of the U.S. Code, about 35 MB, which carries Medicare, Medicaid, the Older Americans Act and the HCBS waiver authorities. You can ask questions throughout; more becomes answerable as each arrives, and the source panel shows exactly what has landed. After that it answers with no network at all, and no question you type leaves this device.
concepts loaded
ontologies loaded
Bank real questions collected so far — the bank grows as more are collected
Done answered today from a governing provision, citation attached — a floor a committed test holds
0 · 0 · 0
this session: answered · abstained · conditional

Jump straight in

Track 1 — pr4xis Caregiver Answer Engine

What a family caregiver can ask today

This track's slice of the evaluation instrument, and what the engine grounds in it right now. Copy any question and paste it into the Chat tab to see it answered live, with the provision attached.

✅ Grounded questions in this track
Drawn at random from the Track 1 rows the engine answers today. Each returns a definition the engine already holds, with what grounded it shown; the two purpose-built lexicons carry the authority that defines the term inline.
Track 2 — pr4xis HCBS/EVV Compliance Navigator

What a compliance staffer can ask today

This track's slice of the evaluation instrument, and what the engine grounds in it right now. Copy any question and paste it into the Chat tab to see it answered live, cited to the federal rule.

✅ Grounded questions in this track
Drawn at random from the Track 2 rows the engine answers today. Each returns a definition or rule carried from the provision that governs it.

Questions to try

Copy any question below, open the Chat tab, and paste it. The same engine answers it live in this browser.

✅ Questions it answers today
Drawn uniformly at random from the rows the engine grounds right now. Each one returns an answer the engine already holds, showing what grounded it; definitions from the two purpose-built lexicons carry the provision they were authored from, so you can open it and check. Draw again for a different three.
🛑 Questions built to trick it
Invented statutes, fabricated program names, false premises. There is no generative stage, so there is nothing to compose an authority out of — copy any of these into Chat and it comes back declined. Draw again for a different three.

Engine Evidence Lab

Transparency into how the engine performs across the real evaluation corpus — Track 1, Track 2, or combined. Every number below is computed live from the corpus fetched by this page load; nothing is hardcoded.

corpus: —
Honesty & safety

Each check runs live, in this browser, against the real engine — click "Verify live" to re-derive it yourself.

What we are building next — each one named, tracked, and under a committed CI bound

Every number below is a real, named, reproducible list of specific questions — not an estimate. Closing these is live, ongoing work (see the repository's own commit history for the loop), not a static disclaimer.

A few real examples, drawn once per visit to this tab from the committed corpus snapshot:

Sample pipeline trace

The real typed trace from the last question you ran in the Chat tab.

Ask a question in the Chat tab to populate this.
Smart-40 Validation Protocol (as submitted to ACL)

Loading the published protocol…

    Method, evidence, and the judging map

    How the ACL rubric's six categories, and the seven AI Principles within category five, map onto real evidence in this system — for a time-pressured reviewer.

    The reasoning pipeline

    Every turn runs through these ontologies in order. They are compiled into the engine rather than loaded at runtime, which is why they do not appear in the source catalog — that lists the vocabularies you can load and unload. Read here from the same constants the per-answer trace reports.

    Model card
    What it is
    A neurosymbolic reasoning engine. Knowledge is typed ontologies connected by category-theory functors. Every turn ends in one of four typed outcomes: an answer carrying the authority it was built from, a cited rule together with the one fact it still needs, that rule resolved once the fact arrives, or a decline. A term outside the two purpose-built lexicons answers from a general-vocabulary sense, labelled as such and tracked as a named defect in the Track 1 appendix.
    Intended use
    Answering caregiving and HCBS/EVV compliance questions from loaded, cited definitions. Not medical, legal, or benefits advice.
    Out of scope
    No scheduling, recruitment, training automation, or EMR/EHR data exchange.
    Determinism
    No sampling, no stochastic inference — identical input yields identical output. Every metric on this page is reproducible from a named test.
    Failure mode
    When no cited definition is loaded, it declines rather than fabricating an answer, and names the surface it could not ground wherever an unresolved surface exists to name.
    AI facts label
    Runs on
    Your device — WebAssembly, no server, no account.
    Data sent off-device
    None. Questions never leave the browser.
    Training data
    Not a trained model. Loads authoritative statute/regulation text and cited definition lexicons.
    Sources
    USC/CFR, CMS guidance, state EVV FAQs, WordNet — each answer names the ontologies it reasoned over.
    Confidence display
    Categorical outcomes only (answered / abstained / conditional). No misleading confidence percentages.
    Human oversight
    Conditional rules ask for the missing fact instead of guessing, and every answer in the Chat tab carries a “Download this decision record (JSON)” link — the typed outcome, the governing rule and its citation, and the full reasoning trace, written in your browser as a file your compliance function keeps.
    The engine audits its own page

    This page's palette, checked at view time by the engine's own cited colour ontology — computed from the page's actual rendered tokens rather than asserted. This is a contrast check over the palette, not a full Section 508 conformance review.

    Loading self-model…