Discover MCPs & agents
Loading MCPs and agents…
Loading MCPs and agents…
Local-first PC activity recorder for Windows. Captures screenshots, accessibility text, UI events, clipboard activity, and optional audio transcription into a local SQLite database — then exposes everything through a REST API, browser UI, and MCP server.
From the repo.
DeskMate is a local-first desktop activity recorder for Windows. It captures your screenshots, on-screen text (OCR / accessibility tree), keyboard/mouse/clipboard events, and optional audio transcription into a SQLite database on your own machine, then lets you use that data through a browser UI, a REST API, and an MCP server that plugs straight into AI coding agents like Claude Code and Cursor — so your assistant can search your activity, ask natural-language questions, and run auto-generated reports (see MCP server).
All your data stays on your own computer — nothing is uploaded to the cloud.
If you just want to try it out, follow these 4 steps. You don't need to understand anything below this section.
You need Python 3.10 or newer.
python --version — you should see Python 3.10.x
or higher.git clone <repo-url>
cd deskmate
# Create an isolated Python environment (keeps your system clean)
python -m venv .venv
.\.venv\Scripts\Activate.ps1
Every time you open a new terminal to use DeskMate,
cdinto the folder and run.\.venv\Scripts\Activate.ps1again to activate the environment. You'll know it worked when you see(.venv)at the start of your prompt.
pip install -e ".[ocr-winrt,audio,vad,mcp]"
This installs: core features + built-in Windows OCR + audio transcription + voice
activity detection + MCP. That's enough for most people.
(What each [...] means and what else you can add → see Install features as needed.)
deskmate ui
Your browser opens at http://127.0.0.1:3030/deskmate and DeskMate starts recording.
On first run it auto-creates a config and data folder under
C:\Users\<your-username>\.deskmate\. You're up and running.
To stop: press
Ctrl + Cin the terminal wheredeskmate uiis running.
DeskMate splits the "heavy" features into optional modules (extras) — install only what you need, not everything at once. The install format is always:
pip install -e ".[module-name]" # install one
pip install -e ".[module1,module2]" # install several (comma-separated, NO spaces)
| Module | What it adds | When you need it / notes |
|---|---|---|
ocr-winrt | Built-in Windows OCR | Recommended default, no extra downloads |
ocr-rapidocr | PP-OCR (via OpenVINO) | Best for Chinese / small UI text, models bundled |
ocr-tesseract | Tesseract OCR | Requires Tesseract installed and on PATH |
audio | Recording + Whisper transcription | Needs a mic or system audio (WASAPI). Default transcription backend |
audio-openvino | Whisper on Intel NPU/GPU/CPU | Accelerated on Intel Core Ultra (NPU) or Arc GPU, see below |
vad | Voice activity detection (Silero VAD) | Skips silent segments to save compute; install alongside audio |
speaker | Speaker diarization | Tells "who said what" |
redact-onnx | PII redaction | Optional privacy-protection model |
semantic | Semantic (vector) search | Search by meaning, not just keywords |
pipes | Scheduled tasks (YAML) | Run reports automatically on a schedule |
mcp | MCP server | Integrate with AI agents / Claude, etc. |
training | Local LoRA fine-tuning | Fine-tune a small model on your own data, see training docs |
full | Most common extras at once | = tesseract + audio + vad + redact + semantic + pipes. Excludes OpenVINO & training |
dev | Dev tools (pytest, ruff) | Only needed to modify code / run tests |
Beginner tip: start with the default combo from Step 3. Add a module later when you decide you want a feature (e.g. Chinese OCR, semantic search) — re-running
pip install -e ".[...]"is safe.Battery Saver / Power Manager is built in and uses Windows APIs through the standard library; it does not need an extra package.
If you have the built .whl (see Building a wheel below)
instead of the source tree, use the same [extras] syntax on the wheel file:
:: core only
pip install deskmate-0.1.0-py3-none-any.whl
:: recommended combo (same as Step 3)
pip install "deskmate-0.1.0-py3-none-any.whl[ocr-winrt,audio,vad,mcp]"
Install (almost) everything in one go — every optional feature at once. Note there is no single "all" extra, so list them explicitly. This pulls in the heavy deps (PyTorch, OpenVINO, ONNX Runtime, Whisper, RapidOCR…), so expect a large download:
pip install "deskmate-0.1.0-py3-none-any.whl[ocr-tesseract,ocr-rapidocr,ocr-winrt,audio,audio-openvino,vad,speaker,redact-onnx,semantic,mcp,pipes]"
This deliberately omits
training(LoRA fine-tuning — only needed if you train models; see training docs),redact-onnx-dml(the DirectML build ofredact-onnx— pick one, not both), anddev(test tooling). The same works on the source checkout — just swap the wheel filename for-e ".", e.g.pip install -e ".[ocr-winrt,audio,vad,speaker,semantic,mcp,pipes]".⚠️ Most extras pull Windows-only / GPU-specific wheels and large model runtimes. Don't install everything unless you actually need it — the recommended combo is enough for most users.
To produce a distributable wheel (e.g. to install on another machine):
pip install build
python -m build --wheel :: output: dist\deskmate-0.1.0-py3-none-any.whl
The wheel bundles all of DeskMate's own code and the built-in apps
(deskmate/apps/), but not third-party dependencies — those are recorded as
requirements and fetched by pip install from PyPI. For an offline install,
pre-download everything into a folder first:
pip download "deskmate-0.1.0-py3-none-any.whl[ocr-winrt,audio,vad,mcp]" -d wheelhouse
:: then on the offline machine:
pip install --no-index --find-links wheelhouse "deskmate[ocr-winrt,audio,vad,mcp]"
training extra (LoRA fine-tuning)The training extra (torch + unsloth + transformers + peft + accelerate)
is deliberately left out of the "everything" command above, because the right
torch build depends on your accelerator. A pip [extra] can't pin a
hardware-specific wheel (the Intel XPU torch lives on PyTorch's own index, not
PyPI). Below, WHL = deskmate-0.1.0-py3-none-any.whl; on a source checkout use
-e "." instead.
Intel Arc / Core-Ultra iGPU (XPU): the XPU torch wheel and an
oneAPI/Level-Zero/MSVC toolchain are required — there are several non-obvious
gotchas, so this path is scripted. Install the extra, then run the setup script:
pip install "WHL[training]"
scripts\setup-intel-xpu.bat :: pulls torch 2.10.0+xpu + Level-Zero SDK headers
Full walkthrough (oneAPI, MSVC, the toolchain gotchas) → training docs.
CPU only (no GPU — slow, for validating the pipeline on a tiny base like
Qwen/Qwen3-0.6B, not real ≥1B runs):
pip install "WHL[training]" :: the default torch is already CPU-only
Verify XPU is live after setup:
python -c "import torch; print(torch.__version__, getattr(torch, 'xpu', None) is not None and torch.xpu.is_available())" :: expect: 2.10.0+xpu True
DeskMate's natural-language Q&A (Ask) and the LLM apps under My Apps (daily recap, todo extraction, meeting summaries…) need a local LLM server, Ollama. If you don't use these, you can skip this section — recording and search work fine without it.
💡 Recommended: use the Ollama OpenVINO build on Intel hardware If you're on an Intel CPU / Arc GPU / Core Ultra (NPU), we strongly recommend Ollama-OV (OpenVINO backend) — it runs models on Intel CPU/GPU/NPU, faster and more power-efficient: 👉 https://github.com/zhaohb/ollama_openvino . On regular GPUs / other platforms, the official Ollama works fine.
Option A — Ollama OpenVINO build (recommended on Intel)
ollama.exe and start the
service (see that repo's README).Modelfile with ModelBackend "OpenVINO" → ollama create <name> -f Modelfile).config.toml (below).Option B — Official Ollama (other platforms / simplest)
http://127.0.0.1:11434).ollama pull qwen3.Final step for both options — in C:\Users\<your-username>\.deskmate\config.toml,
set the model name to the one you actually have:
[ollama]
base = "http://127.0.0.1:11434"
model = "qwen3_8b_ov:v1" # OpenVINO build: the name you imported it as.
# Official build: the name you pulled (e.g. qwen3).
Then restart deskmate ui — you can now use Ask on the Home page and run reports under My Apps.
These are more specialized capabilities - beginners can ignore them for now and come back when needed:
| Feature | One-liner | Docs |
|---|---|---|
| OpenVINO-accelerated transcription | Run Whisper faster on Intel NPU/GPU | docs/04-audio.md |
| Live speech translation | Translate as you speak, shown in the UI | docs/18-live-translation.md |
| Learning sessions | Start/end a study block, recap, review queue, knowledge graph | — |
| Meetings | Hand-declared sessions: transcript, screen text, recap video, summary | docs/09-meeting-workflow.md |
| Gmail / Outlook integration | Search real mailboxes in Ask and apps | docs/11-connections.md |
| Battery Saver | On battery, move background AI work and user-selected apps onto efficient cores; the SPA view and live status text are localized in Chinese/English | docs/22-power-manager.md |
| Local LoRA fine-tuning | Train a small model on your data (incl. Intel iGPU) | docs/16-learning-training.md |
| All technical design docs | Architecture of every module | docs/README.md |
OpenVINO Whisper quick setup (Intel devices):
pip install -e ".[audio-openvino,vad]"
Then in config.toml:
[audio]
whisper_backend = "openvino_genai" # default is "onnx_cpu"
openvino_device = "NPU" # NPU | GPU | CPU | AUTO
The model is auto-downloaded from ModelScope on first use; if the chosen device fails to load, it automatically falls back to CPU.
LoRA training (Intel GPU) has a one-click setup script — see scripts/setup-intel-xpu.bat and docs/16.
ocr-winrt needs nothing extra; ocr-tesseract needs Tesseract on PATH;
ocr-rapidocr just needs the extra (best for Chinese)deskmate ui # record + API + auto-open /deskmate
deskmate ui --no-run-daemon # view existing data only, don't start new recording
| Page | What it does |
|---|---|
| Home | Health status, recent activity, Ask (natural-language queries) |
| Apps | Run built-in and plugin LLM apps; configure schedules and time ranges |
| Connect Gmail / Outlook and review mailbox integration status | |
| Timeline | Browse the screenshot timeline |
| Transcripts | Audio transcriptions |
| Translate | Live speech translation controls and translated transcript stream |
| Events | Keyboard / mouse / clipboard / window-focus events |
| Learning | Start/end a study session; recap, review queue, knowledge graph |
| Todos | Action items extracted from email + meetings |
| Meetings | Start/end a session; audio + screen text, recap video, summary |
| Capture | Capture-control view for recording sources and runtime toggles |
| Training | Local LoRA fine-tuning |
| Model Service | Start/stop local model backends and inspect service logs |
| Diagnostics | Environment self-checks for common setup issues |
| Reminders | Proactive habit reminders, acknowledgement, and history |
| Battery Saver | EcoQoS battery saver for background AI workers and selected apps; UI text is localized in Chinese/English |
| Settings | Config and monitors |
deskmate serve # HTTP API only (records by default)
deskmate record # recorder only
deskmate capture-once # capture one frame manually
deskmate search "query" # keyword search via the API
deskmate mcp # MCP server for AI agents (see "MCP server" below; API must be running)
Split API and recorder into two processes:
deskmate record
deskmate serve --no-run-daemon
Built-in apps live in deskmate/apps/ (and ship inside the installed package);
your own apps go in ~/.deskmate/apps/plugins/. Run them from the My Apps
page or the CLI. Full list, examples, and how to write your own in
deskmate/apps/README.md.
| App | Purpose |
|---|---|
todo-list | Unified checkbox todos from email + meetings |
meeting-summary | Summarize a meeting you started/stopped |
user-learning | Recap a study session (concepts + notes) |
email-digest | Inbox overview |
email-compose | Draft / reply via Gmail or Outlook |
day-recap / time-breakdown / ai-habits | Daily recap / time split / AI-tool usage habits |
user-profile / habit-report | Multi-day user profile / routine (default: last 7 days) |
These apps need both Ollama and the DeskMate API running.
Config file: C:\Users\<your-username>\.deskmate\config.toml (auto-created on first run).
Main sections: [capture], [a11y], [ocr], [audio], [ollama], [redact],
[filters], [server]
Override with environment variables (prefix DESKMATE_):
$env:DESKMATE_SERVER__PORT = "4040"
$env:DESKMATE_AUDIO__ENABLED = "true"
C:\Users\<your-username>\.deskmate\
├── config.toml # configuration
├── data.db # SQLite database (with full-text search)
├── frames\ # screenshots
├── audio\ # audio chunks (when enabled)
├── videos\ # video chunks
├── checkpoints\ # LoRA training artifacts (when training)
├── apps\ # LLM app output reports; apps\plugins\ for your own apps
└── logs\ # logs
To wipe and start over: stop DeskMate, then delete data.db and the folders above.
Default base URL: http://127.0.0.1:3030. Full endpoint list is in the running
server's /docs (OpenAPI). Common ones:
GET /health GET /search?q=... POST /ask
GET /frames POST /capture GET /meetings
GET /todos POST /todos PATCH /todos/{id}
DeskMate ships an MCP stdio server so other AI
agents (Claude Desktop, IDE assistants, custom clients) can query your local
activity and drive DeskMate's Q&A and report apps. It's a thin proxy over the
local HTTP API, so the DeskMate API must already be running (deskmate serve
or deskmate ui).
Install & launch
pip install -e ".[mcp]" # or add `mcp` to your extras combo
deskmate serve # in one terminal — the API on :3030
deskmate mcp # in another — the MCP stdio server
Tools exposed (all prefixed deskmate_ so they don't collide with other MCP servers):
| Tool | Kind | What it does |
|---|---|---|
deskmate_search | read | Full-text search over frames (OCR + a11y text), UI events, audio transcripts |
deskmate_recent_frames | read | List the latest captured frames |
deskmate_recent_events | read | List latest UI events (clicks, focus, key text, clipboard) |
deskmate_health | read | Daemon liveness + counters |
deskmate_capture_once | action | Trigger a paired capture now |
deskmate_ask | action | Natural-language Q&A: an LLM searches your local context and answers (slow) |
deskmate_list_apps | read | List the report apps (day-recap, standup-update, todo-list, meeting-summary, email-compose, …) |
deskmate_run_app | action | Run a report app and return its result (slow; params vary per app) |
deskmate_learning_sessions | read | List recent study sessions |
deskmate_learning_reviews | read | Due review cards (spaced repetition) |
deskmate_list_app_outputs | read | List an app's past run outputs |
deskmate_get_app_output | read | Fetch one output file (markdown/json) of a past run |
deskmate_askanddeskmate_run_appinvoke a local LLM, so they need Ollama + the API running and can take tens of seconds to minutes.
Once DeskMate is registered, your coding assistant knows what you've been doing locally — no copy-pasting context. While you work in Claude Code or Cursor you can ask:
deskmate_ask — natural-language Q&A over your activity
deskmate_search — full-text search over OCR / UI / audio
deskmate_list_apps / run_app — report apps
deskmate_learning_sessions / reviews — study recap
deskmate_recent_events / recent_frames — raw activity stream
deskmate_health / capture_once — status & instant capture
deskmate_list_app_outputs / get_app_output — past reports
The assistant reads your local activity record on demand and answers in-context, entirely on your machine.
Point command at the Python env that has the mcp extra installed.
Claude Code (CLI):
claude mcp add deskmate -e DESKMATE_API=http://127.0.0.1:3030 \
-- C:/path/to/env/python.exe -m deskmate mcp
# verify: claude mcp list (should show deskmate ✓ connected)
Cursor — add to ~/.cursor/mcp.json (global) or <project>/.cursor/mcp.json:
{
"mcpServers": {
"deskmate": {
"command": "C:/path/to/env/python.exe",
"args": ["-m", "deskmate", "mcp"],
"env": { "DESKMATE_API": "http://127.0.0.1:3030" }
}
}
}
Claude Desktop — same block in claude_desktop_config.json.
After editing config, fully restart the client — MCP servers are only loaded at startup.
DESKMATE_API (optional) overrides the API base URL; defaults to
http://127.0.0.1:3030.deskmate_ask and deskmate_run_app invoke a local
LLM and can take minutes. The server-side budget defaults to 30 min and is
configurable via DESKMATE_MCP_ASK_TIMEOUT / DESKMATE_MCP_RUN_APP_TIMEOUT
(seconds, in the server's env). These tools also stream MCP progress every
~15s, which keeps most clients from cutting the call off; if your client still
times out, raise its tool timeout with MCP_TOOL_TIMEOUT (milliseconds; e.g.
1800000 for 30 min).HTTP(S)_PROXY, so 127.0.0.1 traffic
is never routed through a proxy.mcp extra.ollama list should list models) and that [ollama] model in config.toml matches
a model you actually have. On Intel hardware, use the
Ollama OpenVINO build (the default model
name qwen3_8b_ov:v1 is meant for it).deskmate not found in a new terminal? You didn't activate the virtual
environment. cd into the project folder, then run .\.venv\Scripts\Activate.ps1.MIT — see LICENSE.
DeskMate is free to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, subject to the MIT license terms.
Replace {MCP_ENDPOINT_URL} with this MCP’s endpoint URL (from its repo or docs above). No API key — you connect directly.
Tool
OS
Config file: ~/.cursor/mcp.json
{
"mcpServers": {
"mcp-server": {
"url": "{MCP_ENDPOINT_URL}"
}
}
}Paste into mcpServers in the config file. Restart Cursor after saving.
If this MCP is also published on mcpchannel.ai, you can subscribe from Browse and use the gateway config there instead.