Discover MCPs & agents
Loading MCPs and agents…
Loading MCPs and agents…
Stability-first operations CLI for long-running AI agent workspaces.
From the repo.
The local operations layer for long-running AI agents.
Coding, ops, research, automation — any agent that runs for hours on the same workspace.
Profiles before commands · Checkpoints before risky work · Durable history after the chat is gone.
Landing · Quickstart · What Helm does · Workflows · Docs · 한국어
pip install helm-agent-ops
helm init --path ~/.helm/workspace
export HELM_WORKSPACE=~/.helm/workspace
Run your first inspection under a declared risk profile:
helm profile run inspect_local --task-name "first look" -- git status --short
helm status --brief
helm dashboard
The first command produces a guarded execution record. The second shows what just happened in plain English. The third lays out the workspace state on one page.
No PyPI? Use the bootstrap installer:
curl -fsSL https://raw.githubusercontent.com/JDeun/Helm/main/install.sh | bash
Long-running AI agents drift. They forget prior decisions, execute risky actions before you can stop them, and leave behind a chat log nobody can audit a week later — regardless of whether the agent is editing code, running ops, organizing notes, browsing sites, or chaining tool calls.
Helm is a thin, file-backed operations layer that sits around your existing agent runtime. It does not replace your agent. It makes the agent's work boundable, recoverable, and reviewable.
The model proposes actions; the harness validates, authorizes, executes, records, and returns observations. Safety and completion claims should come from execution evidence, not from prompt advice or a compacted chat transcript.
| Without Helm | With Helm |
|---|---|
| Risky commands run as soon as the agent decides | Commands run under a declared execution profile with a guard check |
| Multi-step or multi-file changes leave you guessing what happened | Checkpoint created before the work; visible rollback point |
| "What did the agent do yesterday?" → scroll the chat | Local task ledger, command log, dashboard, markdown report |
| Context lives in the chat window | File-backed memory + ranked retrieval rehydrates the next session |
| Skill rules live in prompts | SKILL.md + contract.json enforce policy at run time |
If your agent only runs one-off demos, you do not need Helm. If you run it for hours on the same workspace — coding, ops, knowledge capture, or any mix — you do.
🛡️ Guard before execution
|
💾 Recover after the fact
|
🧭 Operate over time
|

helm profile run inspect_local --task-name "inspect current repository" -- git status --short
helm checkpoint create --label before-risky-work --include $HELM_WORKSPACE
helm report --format markdown
helm dashboard
Each command leaves a structured record on disk: task ledger, command log, checkpoint record, dashboard summary. None of it requires the agent to remember anything.
helm doctor
helm status --brief
helm dashboard
helm reconcile # dry-run: report reference drift vs the desired snapshot
helm verify-contract # assert behavioral operating invariants still hold
helm profile run inspect_local --task-name "inspect repository state" -- git status --short
helm profile run workspace_edit --task-name "tighten typing in api/" -- ruff check api/
helm survey
helm onboard --use-detected --dry-run
helm onboard --use-detected
helm checkpoint-recommend
helm checkpoint list
helm task list --status running
helm task doctor
helm report --format markdown
helm context --mode decisions --explain-ranking --json
helm context --mode timeline --since 2026-05-01
helm context --mode entity --entity project_helm
helm context --mode reflect-candidates
helm privacy scan --text "Contact alice@example.com" --json
helm privacy tokenize --scope task-123 --text "Contact alice@example.com"
helm skill-lifecycle negative-claims --persist
helm skill-lifecycle revalidation-due
helm skill-lifecycle revalidate-claim \
--skill old-skill \
--claim-id sha256:abc123 \
--status resolved \
--note "command now exists"
helm run-contract --json
helm capability-diff --json
helm skill-promotion digest --json
helm shadow-report --since 14 --format md --with-recommendations
helm health state --json
helm health select --profile inspect_local --context-tokens 4096 --json
python3 scripts/model_health_probe.py probe --model omfm/balanced --json
Every command also accepts
--path /custom/workspaceif you do not want to use$HELM_WORKSPACE. The demo workspace atexamples/demo-workspaceis safe to point at.
Current release: v1.0.0 — released 2026-09-28.
Breaking change. The four unprefixed top-level packages scripts, commands,
references and memory_tree now live under the helm package: helm.scripts,
helm.commands, helm.references, helm.memory_tree. Installing helm-agent-ops
previously claimed scripts and commands — two of the most common directory names
in Python projects — and could silently shadow a user's own package of that name;
that happened to helm's own only user. There is no compatibility alias —
replace from scripts.X import Y with from helm.scripts.X import Y, and likewise
for commands, references and memory_tree, before upgrading.
helm, helm_context, helm_frontmatter,
helm_state_model, helm_workspace) instead of nine.helm.py is now a package (helm/cli.py plus a lazy helm/__init__.py). The
helm console script and import helm; helm.main([...]) are unchanged;
python -m helm is now also supported.v0.13.0 — released 2026-07-16. This release imports patterns proven in live agent operation.
helm reconcile re-applies workspace reference files against the packaged desired snapshot — drift-tolerant and idempotent, reporting drift instead of clobbering local overrides.helm verify-contract asserts behavioral operating invariants (guard deny/fail-closed, approval TTL/consume-once, atomic ledger), complementing structural doctor/validate.scripts/skill_router.py), a generic tool/MCP adapter registry (scripts/tool_adapter.py + references/connectors.json), and grounding-by-guidance-injection with a deterministic template fallback (scripts/grounding.py).run_checkpoint uses sys.executable; non-file checkpoint backends fingerprint dependencies before a runtime bump.v0.12.0 — released 2026-07-13. This release makes workflow boundaries and completion claims explicit and inspectable.
references/workflow_units.yaml defines allowed inputs, live sources, mutation surfaces, verification, reporting, handoff, and stop contracts.python3 scripts/workflow_registry.py validates those contracts before adoption.Released 2026-06-24. This patch adds read-only loop validation and conservative external skill-intake classification.
helm loops validate and helm loops inspect validate reusable workflow contracts.helm skill-intake classify and helm skill-intake validate provide a conservative review path for external skill candidates.Released 2026-06-20. This patch keeps task-ledger attribution inspectable across profiled runs and chat memory captures.
experience_attribution.helm memory capture-chat keeps queued / running rows free of final-only memory and attribution payloads.conversation as the selected tool for attribution.Released 2026-05-22. Everything new ships in shadow mode by default — decisions are logged but not enforced until you opt in.
{component, tool, profile, error_class, target, fingerprint} so the same failure is recognizable across runs.OPENCLAW_PAUSE_GATE.allow_single_session, block_mutation, require_user_login, require_confirmation, pause_profile, require_cleanup_evidence) with a runner-side enforcement gate.HELM_MODEL_REPAIR and HELM_SYNTHETIC_RESPOND.helm shadow-report --since 14 --with-recommendations aggregates 14 days of signals and emits ready_to_enforce / needs_more_data / caution / no_signal per feature.See the full v0.10.0 notes and the 13-document docs/harness-engineering/ directory for the design.
Helm runs in a dedicated workspace, treating existing systems as read-only context sources first.
.helm/ inside the workspace.| Category | Better for | Helm adds |
|---|---|---|
| Agent frameworks (LangChain, AutoGen, etc.) | prompts, planners, tool loops, agent graphs | profiles, guard decisions, checkpoints, task ledgers |
| Observability (Langfuse, Helicone, etc.) | hosted traces, service metrics | pre-execution policy + local recovery state |
| Evaluation (DeepEval, Phoenix, etc.) | scoring model output | operational history around repeated human-agent work |
| Shell wrappers (cmd helpers) | command convenience | workspace state, memory capture, reports, recovery discipline |
See deeper comparisons in docs/comparisons/.
| Get started | Core concepts | Advanced |
|---|---|---|
Helm's design follows the findings in Harness Design Determines Operational Stability in Small Language Models, which experimentally studies how planning, verification, and recovery harnesses affect operational stability. Its adaptive-harness direction is also informed by It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers, which shows that harness strictness should be selected by model type and failure mode rather than applied uniformly.
Cite Helm:
@software{helm_2026,
title = {Helm: A stability-first operations layer for long-lived agent workspaces},
author = {Cho, Yong Eun},
year = {2026},
url = {https://github.com/JDeun/Helm},
version = {0.13.0}
}
See CITATION.cff for the machine-readable form.
Issues and pull requests welcome.
CONTRIBUTING.md before opening a PR.python -m pytest -q (currently 1,596 tests).python scripts/release_version_check.py --version <next>.SECURITY.md.scripts/commands/references/memory_tree moved under helm.*, no compat shimCHANGELOG.md · older release notesHelm ships only the public operations layer. It does not include:
The repository is safe to fork, clone, and inspect.
Replace {MCP_ENDPOINT_URL} with this MCP’s endpoint URL (from its repo or docs above). No API key — you connect directly.
Tool
OS
Config file: ~/.cursor/mcp.json
{
"mcpServers": {
"mcp-server": {
"url": "{MCP_ENDPOINT_URL}"
}
}
}Paste into mcpServers in the config file. Restart Cursor after saving.
If this MCP is also published on mcpchannel.ai, you can subscribe from Browse and use the gateway config there instead.