HostleloBlogExplore hosting

Hikmah Stack Review:Give AI Agents Memory You Can Inspect

AI agents can lose corrections, repeat stale claims, and say a task is done before tests run. Hikmah Stack makes memory, disagreement, and uncertainty inspectable through a local Rust kernel. Here is how it works, how to try it, and where its limits remain.

Illustration of memory records, conflicting claims, and a linked ledger for Hikmah Stack
On this page

The short answer

An open-source Rust kernel for traceable memory, visible contradictions, bounded decisions, and more honest completion checks.

Before you start

Git and Rust 1.89 or later for the pinned local CLI example. Cargo downloads build dependencies. Use a fresh non-sensitive demo store; memory is stored in plaintext. The no-default-features build excludes the optional network adapter. Review host permissions and provider data destinations before integration.

Disclosure: This creator-led spotlight features Juber Shaikh's own open-source project. It combines public-source review with the bounded checks described below, rather than an independent customer endorsement.

An open-source Rust kernel for traceable memory, visible contradictions, bounded decisions, and more honest completion checks.

AI agents can produce a convincing answer while losing track of an earlier correction, treating a guess as evidence, or declaring a task complete before the tests run. These problems become more expensive when the same agent works across sessions. A chat transcript alone gives developers little structure for deciding which claims deserve to survive.

Hikmah Stack tackles that engineering problem with explicit state and deterministic controls. Its most interesting idea is to make memory, uncertainty, and accountability inspectable in ordinary software. Developers can examine a stored claim, see its source, follow a correction, and ask why it was retrieved.

This review covers the public Hikmah Stack source at commit 2721bc9, reviewed on October 3, 2026, with bounded local validation in a fresh Linux workspace. The package identifies itself as version 3.1.0. Findings below distinguish source inspection, executed checks, repository-published measurements, and integration work a developer would still need to do.

What is Hikmah Stack?

Created and maintained by Juber Shaikh, Hikmah Stack combines a Rust library and command-line application with portable instruction skills. The runtime stores typed memories, tracks structured disagreements, retrieves relevant traces, manages commitments, and evaluates bounded planning or decision inputs.

The separation matters. The Rust code enforces rules over the data it receives. The skills guide how an AI host approaches judgment, debugging, decisions, and delivery. Installing the skills does not automatically connect every conversation to the memory kernel. Your host or application still has to call the CLI or library at the right points.

The package uses the MIT license, with its attribution and license-notice requirements. Rust developers can inspect, adapt, and integrate the implementation; application developers can start with its JSON-producing CLI. The package declares Rust 1.89 as its minimum supported version. Package manifest, license.

Why an agent needs more than a growing conversation

Consider a coding assistant investigating failed deployments. On Monday, a runbook says the database uses port 5432. On Tuesday, migration notes say it moved to 5433. By Wednesday, a plain text summary may contain only the most recent sentence, with no indication that another source disagrees.

Hikmah provides a more useful representation. Both statements can remain as separate traces with a common claim key, different values, and distinct sources. The conflict becomes something the application can display and resolve deliberately.

That makes it especially interesting for incident investigations, long-running coding work, and experiments where developers need to reconstruct what their agent knew. The advantage is concrete: fewer important distinctions disappear into a rewritten summary. Whether that improves task success still needs measurement in the application using it.

Inside a memory trace

The central data structure contains content, a type, tags, timestamps, confidence, salience, privacy classification, and provenance. It can also carry a deadline, a structured claim, or a link to the trace it supersedes. Supported types include observations, beliefs, procedures, commitments, corrections, predictions, and outcomes.

Those types give downstream code something meaningful to work with. A forecast can stay separate from an observed result. A commitment can carry a deadline. A correction can replace an active belief without erasing its history.

Validation rejects empty content, invalid numeric ranges, and incomplete claim pairs. Model-labeled traces cannot be marked verified. Detected AI-session writes receive a session locator and cannot claim human verification. These checks are useful guardrails, although the source labels remain caller assertions and session detection uses environment variables. They do not authenticate a person or make malicious local code trustworthy. Trace definitions and validation, session boundary.

How memory moves through the system

A typical integration has a small, understandable loop:

  1. The host supplies a proposed memory with its source and metadata
  2. The kernel validates it before writing
  3. The local ledger records the event and its place in the hash chain
  4. Recall selects relevant active traces and exposes conflicts or correction links
  5. New observations, fulfilled commitments, and outcomes extend the history
Hikmah architecture showing explicit host calls into a local Rust kernel and JSONL ledger, with a separate optional Jev network boundary

Original explanatory diagram of the reviewed implementation. The host owns scheduling and external actions.

The default store is .hikmah/memory.jsonl. Its implementation uses sequence numbers, BLAKE3 hashes, an exclusive writer lock, and a sidecar head file. The head records the latest acknowledged position. Torn final records can be handled without discarding acknowledged history, and ordinary writes refuse known ledger/head inconsistencies. Ledger implementation.

This is useful audit infrastructure, with an important boundary: the hashes are unkeyed. Someone able to rewrite both the ledger and its head file can recompute them. An externally protected expected head adds a stronger check. Successful verification establishes the applicable integrity checks, not the truth of every stored statement.

Try the local memory workflow

The basic demo needs Git, a suitable Rust toolchain, and no model account. Cargo may download dependencies during the first build. Using the no-default-features build excludes the optional network adapter from the resulting binary.

git clone https://github.com/CodeWithJuber/hikmah-stack.git
cd hikmah-stack
git checkout 2721bc965043c5b03fe60c0a0c86983762962de7
cargo build -p hikmah-kernel --no-default-features

./target/debug/hikmah init --store .hikmah/review-demo.jsonl
./target/debug/hikmah remember --store .hikmah/review-demo.jsonl \
  --kind observation --source demo:incident \
  --content "Evening deploys fail when database locks time out" \
  --tag deploy --confidence 0.9
./target/debug/hikmah recall --store .hikmah/review-demo.jsonl \
  --query "database deploys fail"

Use a fresh demo store so previous experiments do not change the results. The JSON response makes a useful inspection point: look for the trace identifier, provenance, score, and separate retrieval channels. The confidence entered above is an assertion supplied to the program, not evidence that the incident has been independently verified.

Now introduce a disagreement:

./target/debug/hikmah remember --store .hikmah/review-demo.jsonl \
  --kind belief --source demo:runbook \
  --content "Database port is 5432" \
  --claim-key db.primary.port --claim-value 5432
./target/debug/hikmah remember --store .hikmah/review-demo.jsonl \
  --kind belief --source demo:migration \
  --content "Database port is 5433" \
  --claim-key db.primary.port --claim-value 5433
./target/debug/hikmah conflicts --store .hikmah/review-demo.jsonl
./target/debug/hikmah verify-ledger --store .hikmah/review-demo.jsonl

The conflict command exposes the competing values. It does not choose which source is correct. That decision needs evidence from your environment. Claim keys are normalized, while values preserve meaningful case differences such as filesystem paths. Conflict implementation.

Repository-supplied terminal recording showing Hikmah memory creation, recall, and ledger verification

Repository-supplied recording of the v3.1.0 CLI. This existing demo illustrates the memory flow; it is separate from the commands above.

The review's fresh Linux checks passed the memory, conflict, and ledger-verification demo. These checks exercise the fictional local workflow; the animation above remains the maintainer's earlier recording.

Recall that explains its priorities

Hikmah's recall is lexical. It tokenizes text, applies light English stemming, supports CJK character bigrams, and uses term and tag overlap. Recency, salience, confidence, provenance, and urgency can adjust the score after relevance is established. A highly confident unrelated memory cannot qualify simply because its metadata looks impressive.

Results include their scoring channels, unresolved conflicts, and supersession links. Near-duplicate memories can be folded together. Predictions are excluded by default, and superseded traces require an explicit historical lookup. Overdue commitments receive special handling so they can surface without matching the query's vocabulary when recall runs; this is not a background reminder service. Recall code.

This design is easy to inspect and debug. It also has a clear tradeoff: different wording can hide a related idea. The published September 26 retrieval evaluation placed Hikmah below its BM25 baseline on all three measured datasets or samples. For example, SciFact NDCG@10 was 0.5358 for Hikmah and 0.6617 for BM25. These were document-retrieval experiments, including only 100 of 648 FiQA queries, rather than proof of agent-memory effectiveness. Published retrieval results and limitations.

Turn follow-through into inspectable state

Commitments extend the same approach to unfinished work. Record a commitment with a deadline, inspect what is due, and mark it fulfilled when the work actually happens. For an incident assistant, that could mean retaining the obligation to write a postmortem after the immediate troubleshooting conversation ends.

./target/debug/hikmah remember --store .hikmah/review-demo.jsonl \
  --kind commitment --source demo:incident \
  --content "Write the deployment incident summary" --deadline +48h
./target/debug/hikmah commitments --store .hikmah/review-demo.jsonl \
  --within-hours 168

Your host still owns scheduling and notifications. The kernel provides queryable state that a scheduler can inspect. Commitment query implementation.

Consolidation is similarly explicit: it groups compatible structured claims and emits proposals with supporting trace IDs, claimed sources, conflicts, and eligibility. It does not automatically promote a generated summary into established knowledge. Model-authored traces and predictions cannot supply supporting evidence, while source independence is assessed from normalized source names rather than authenticated identities. Consolidation implementation.

Decisions can keep uncertainty visible

The decision evaluator accepts options, weighted criteria, evidence-backed scores, and hard blocks. Missing scores produce intervals representing the unresolved portion of the evaluation. Options rank using their lower bounds, with a documented preference for reversible choices in certain close, weak-evidence situations.

An especially useful detail is the separation between model estimates and evidence. An estimate may influence ranking, but it does not increase evidence coverage or make a decision decisive. A result becomes decisive only when the recommended option's evidence lower bound exceeds every competing admissible option's evidence upper bound. All of this remains conditional on the criteria and inputs supplied by the caller. Decision evaluator.

The repository also includes bounded symbolic planning and five deterministic challenge lanes covering evidence, memory, risk, human impact, and delivery. These lanes evaluate supplied counts. Their presence does not mean multiple autonomous agents independently investigated the problem.

./target/debug/hikmah plan --problem examples/plan-problem.json
./target/debug/hikmah decide --frame examples/decision-frame.json
./target/debug/hikmah ask --request examples/decision-request.json

The final command uses the default no-engine path and abstains. That is a useful integration test: your application should handle an explicit unknown rather than inventing a result. Planner, challenge lanes.

Optional model decisions and measured forecasts

The typed decision port accepts bounded choice, score, and boolean-probability questions. Its admission logic rejects an entire engine response when an answer violates the request contract. The optional TypeSafe Jev adapter implements that interface; the text proposal interface currently supplies only a no-model implementation.

Engine answers can be recorded as unverified predictions and later paired with observed outcomes. Calibration reports include Brier scores and expected calibration error, grouped by forecaster and question family. The software requires at least 50 scored outcomes plus additional statistical conditions before assigning its calibrated label. Even that label describes the measured family, not a guarantee about future tasks. Typed decision documentation, calibration implementation.

Jev requires a build with its feature enabled, explicit engine selection, and an API key. Selected request content then leaves the machine. The completion hook's engine mode sends up to 8,000 characters of the last assistant message. Review the destination and data before enabling it. The local demo above does not need it. Adapter, network boundary.

What Truth Gate can actually catch

Truth Gate examines completion messages for a narrow contradiction: a response claims the work is complete while also acknowledging unfinished or deferred work. Its rules account for negation, quoted code, and several innocent uses of words such as TODO or placeholder.

You can inspect the local rules without connecting a provider:

printf '%s' '{"last_assistant_message":"Done. TODO: tests"}' \
  | ./target/debug/hikmah gate-explain

The repository's historical engine evaluation reported 18.2% false-completion recall with a 7.2% false-block rate on its second held-out test split. Those figures describe one agent/model/scaffold setting. The rules were subsequently narrowed without rerunning that benchmark. Treat the numbers as scoped historical evidence, not current universal detection rates. Real test results still matter far more than whether a final message passes a text screen. Evidence notes.

Where it fits in a developer's workflow

The repository packages six skills: Operator Core, Agent Radar, Decision Forge, Ship Guard, Hikmah Orchestrator, and Cognitive Kernel. It includes manifests for Codex, Claude Code, and Kimi, plus documented OpenClaw bundle compatibility.

Review the host-specific instructions before installation. Claude Code and Codex have completion-hook arrangements; Kimi supplies skill routing. OpenClaw's documented integration loads the skills but does not execute the existing completion hooks or expose the Rust kernel as tools. Compatibility is therefore a set of specific adapters, not identical behavior across hosts. Compatibility guide.

For a first adoption, choose one narrow workflow: recording incident observations, preserving corrections during a coding task, or tracking predictions against outcomes. Define what improvement would count, keep a baseline, and evaluate actual work. That will teach you more than adding every capability at once.

Security, maturity, and the honest verdict

The default store is plaintext. Credential-pattern checks and refusal of sensitive persistence are useful protections, but they do not provide encryption, comprehensive data-loss prevention, or access control. Purging a trace hides it from recall while retaining content in the append-only ledger. That matters wherever actual deletion is required.

The project also does not ship an HTTP API, MCP server, embedding index, general model orchestration platform, or enterprise deployment system. At the pinned commit, the review's fresh Linux validation passed 160 Rust tests, with two explicitly ignored. Formatting, Clippy with default and no-default features, package validation, and Python checks also passed. No API or model calls were made, and these checks do not establish live provider behavior, end-to-end host integration, or production reliability.

Separately, the reviewed commit's public validation workflow passed, covering Rust checks, package validation, and hook tests across Linux and Windows jobs. The current local run and the repository's older benchmark reports answer different questions; the benchmark figures above were not reproduced for this review. Security policy, validation workflow.

Hikmah Stack earns attention through the specificity of its engineering. It makes provenance, disagreement, uncertainty, and completion discipline visible enough to inspect and test. The strongest reason to try it is the opportunity to build an agent workflow whose durable state you can explain.

Start with the local demo, inspect the JSON, and challenge a claim. If that approach fits your work, explore the source, review the tests, and contribute a reproducible failure case or integration improvement. That is a practical route toward more accountable AI-assisted software.

Visual credits

The cover is an AI-generated editorial illustration, not application UI. The architecture diagram is an original source-based explanation. The GIF is the repository-supplied v3.1.0 recording, separate from this review's local checks. Demo source; Hikmah Stack, Juber Shaikh, MIT license.

Sources & further reading

  1. Hikmah Stack source
  2. Reviewed source snapshot

What changed

Reviewed Hikmah Stack v3.1.0 at 2721bc965043c5b03fe60c0a0c86983762962de7. Fresh Linux checks: 160 Rust tests passed, 2 ignored, plus formatting, Clippy, package/Python checks and fictional local demos. No live provider calls or all-host integration testing; historical benchmarks were not reproduced.

Originally published . About our editorial updates.

Your next project deserves a better foundation.

Explore hosting built for your next chapter.

Explore hosting