Discover MCPs & agents
Loading MCPs and agents…
Loading MCPs and agents…
Desktop AI agent for Windows — the execution layer for AI: an architecture that expands what any model can do. Files, shell, browser, documents, and a Code door with language servers, Git and pull-request review. OCR, offline speech, real .xlsx/.pptx/.docx output, MCP, a delegated team of agents. Runs on your machine, with any provider.
From the repo.
A Windows desktop app that finishes the work on your machine — files, browser, shell, documents.
ภาษาไทย · Website · Microsoft Store · Download · Community
The source in this repository stops at v1.7.0 (2026-09-15). The product does not. Every release after that still lands here — tags, installers, the portable zip, the Linux engine, checksums and release notes — and the in-app update check, scoop and the Microsoft Store keep working exactly as before.
Why 1.7.0 and not 1.8.0. The source was open through v1.8.0 for one day (20–21 September 2026) and was then rolled back. What arrived between 1.7.0 and 1.8.0 (222 commits in five days) — MCP tools that became a shelf instead of a tool block resent every turn, an editor with a language server behind it and no extension market, a session context written every turn and handed back on reopen, a window nearly half as heavy at rest — is the architectural layer that makes Aetox as cheap to run as it is. From the outside it looks like any other agent app; the inside is where the author spent the most design time, and is the part he chose to keep. Not closed because anyone did anything, but because it is the most valuable part of the work, and he is the one person still building it.
The code you can read here is the real v1.7.0, unchanged, and remains readable for learning under the LICENSE. The measured numbers are still published in full under docs/reports. Bug reports still go to Issues; a star still tells the author something.
Aetox is the execution layer for AI — models provide intelligence, Aetox provides capability.
MODEL ← knows what should be done
↓
┌─────────┐
│ AETOX │ ← eyes, ears, hands, tools, permission
│Execution│
│ Layer │
└────┬────┘
↓
┌─────────────┼─────────────┐
↓ ↓ ↓
Files Browser Shell
↓ ↓ ↓
documents websites programs
The north star of this project. Do not build an AI that answers more. Build an AI that does more.
Aetox is a desktop application for Windows that runs an AI agent against your own machine. You describe what needs doing; it reads and writes real files, runs real commands in a real shell, and drives a real browser you can watch.
It is two self-contained executables, 83.4 MB together — aetox.exe, the window, and since 1.6.0
aetox-engine.exe, the half that thinks and works, beside it. There is no runtime to install
alongside them, no node_modules, no bundled copy of Chromium. It talks to whichever model you point it at —
a hosted API, a subscription you already pay for, or a 9B/35B running in LM Studio or Ollama on
your own GPU (your data never leaves your machine or country — hook it up to Ollama and not a single byte goes anywhere) — and the capability comes from the app rather than from the model's parameters.
That is why a small local model can still read a picture, transcribe a recording, and hand you a
deck that opens in PowerPoint: image_ocr, audio_transcribe and the slide exporter are the
app's, not the model's.
There is a word for that arrangement. Aetox is a harness: the program around a model that gives it tools, a loop, and somewhere to work. The model is the engine; the harness is the car. That is why the same model can be a different product in two apps — and why a score belongs to a model and a harness together rather than to the model alone.
Two concrete jobs, to make that less abstract. "Go through this folder of receipts and give me
one spreadsheet" — it OCRs each image, works out the totals in a JavaScript interpreter
compiled into the binary and shows you the script beside the answer, then writes a real .xlsx
with live formulas. "Find why the login test is flaky" — it greps the repo, runs the suite in
a terminal tab you can watch, reads the failure, and edits the file.
The interface ships in Thai and English, and Thai is the default. Every string exists in both; the language switch is in Settings and in the first-run wizard. This README is in English; ภาษาไทย is here.
.html file that is yours, editable by hand, and
openable on any machine with a browser. Exports as .pdf, .png or .jpg.git, grep and glob over the whole tree, and language
servers the app installs itself.aetox-engine does the work beside it, over one socket. Add a Linux
host under ตั้งค่า › เครื่องระยะไกล and connect with one click: the app ships the engine over
your own ssh, starts it, opens the tunnel and looks after it, and the chat, the files, the
terminal and Git are that machine's. Your provider keys never leave the machine the window
runs on — the engine asks the window to sign each request.Open Yours → Connections, then choose Telegram or Discord:
@BotFather and connect its bot token./pair 123456 command after connecting. Send it from the Telegram chat
or Discord channel you want to authorize (mention the bot when pairing in a Discord server).This is another doorway to Aetox's real main assistant, not a separate bot assistant: it uses the
Assistant desk's identity and tools, and a new conversation starts with the current default model. Each paired Telegram chat or Discord
channel keeps its own conversation history instead of appending to the chat currently open in the
desktop window. The bot accepts messages only from the paired conversation; every member of that
conversation can use it. Mention it in Discord server channels; DMs work directly. Use /new to
start with fresh context. Tokens and pairing data are encrypted at rest, and listeners run only
while Aetox is running.
Windows 10 or later, x64. You do not need an API key to start — a built-in aetox provider
ships five test models that exercise the real machinery (real tool calls, a real delegation to a
sub-agent, a long reasoning stream), so you can see what the app does before signing up for
anything.
Microsoft Store — the one channel with nothing to click past. Microsoft signs the package, so there is no SmartScreen prompt and no antivirus warning, and Windows keeps it updated afterwards. One line, no web page in the way:
winget install --id=9N4KKBRRSCZZ --source=msstore
(To look before installing: winget search aetox. The Store source only answers an --id lookup
with --exact, so winget search --id 9N4KKBRRSCZZ --source msstore comes back empty — a winget
quirk, not a missing listing.)
Prefer to click? apps.microsoft.com/detail/9N4KKBRRSCZZ,
or paste ms-windows-store://pdp/?productid=9N4KKBRRSCZZ into Run (Win+R) to open the Store app
straight away without the web page.
Installer — aetox-amd64-installer.exe (34.6 MB)
Installs into Program Files with a Start menu entry. It carries its own files and nothing else: Tesseract, poppler, ffmpeg and the speech model are fetched later by the app itself, and only for a capability you tick.
Scoop
scoop install https://raw.githubusercontent.com/Mikedev115/Aetox/main/scoop/aetox.json
Portable — the zip,
unpack, run aetox.exe. Since 1.6.0 the zip holds two files — aetox.exe and aetox-engine.exe
— and they stay together. This is the only channel that can update itself in place.
In the terminal — aetox in any folder opens the Code desk full-screen: type the task, watch the
tool calls land above the input, answer an approval card in place. A separate download, not part of
the app above, so the installer and the Store package stay exactly as they are. The console is not
in the Microsoft Store.
aetox-cli-setup.exe
— installs "Aetox CLI" into your user folder (no administrator rights) and puts it on your user
PATH; open a new terminal and type aetox. Remove it from Settings → Apps → Aetox CLI. Or with
scoop:
scoop install https://raw.githubusercontent.com/Mikedev115/Aetox/main/scoop/aetox-cli.json
or aetox-cli-windows-amd64.zip
— unpack anywhere and run .\aetox.exe path add there once to put that folder on your PATH.
When the app is installed on the same machine the
terminal uses the app's engine and shares its keys, chats and memory; without the app it runs the
engine that comes in the zip. aetox chat "task" is one turn with the answer on stdout, for
scripts; aetox --plain is the line-by-line console; /help inside lists every command.
Pick one channel and stay on it. Windows gives a packaged app its own data folder, so a Store install and an installer install are two separate Aetoxes on one machine, with separate settings, history, memory and keys. Installing both is the quickest way to wonder where your chats went.
None of this happens on the Store build — Microsoft signs that one. What follows is about the installer and the zip: two different warnings, two different causes.
"Windows protected your PC", unknown publisher. The installer is not code-signed yet, so Windows has no publisher name to show for it — More info → Run anyway.
"Virus detected", or a name ending in !ml such as Program:Win32/Wacapew.C!ml. A cloud
machine-learning verdict, not a signature: nobody analysed this file and judged it dangerous. The
!ml ending says so. It fires on what the file is rather than what it does — an unsigned binary
whose hash the world has never seen before, which every release is by definition. Desktop apps
built with Go and Wails hit this across the ecosystem; an empty Wails app with no code in it at all
is reported as the same detection.
The installer itself no longer fetches anything third-party. It did until v1.5.7, and that is what earned the original verdict; the installer script carries the whole story. Tesseract, poppler, ffmpeg and the speech model are now downloaded by the app, only for a capability you tick, each pinned to an immutable release tag and verified against a SHA256 compiled into the binary before it is used — a mismatch skips that component rather than proceeding.
A verdict cleared with Microsoft applies to one file, and the next release is a different file, so it can come back until code signing exists. The portable zip is the way past it in the meantime.
The app opens but every provider list is empty, and the engine card says ไม่พบ aetox-engine.exe.
The same verdict, aimed at the second file in the install folder: on 2026-09-15 Defender's cloud
model quarantined aetox-engine.exe from v1.7.1 as Trojan:Script/Wacatac.C!ml, five hours into a
session, on a file that had not changed (Trojan:Script/… is the family Defender uses for an
unsigned executable that starts shells — the engine does, on your behalf, which is its job). The app
cannot answer anything without its engine, so the lists go blank. Open Windows Security →
Protection history, find the entry, Restore and then Allow on device — Restore alone puts
the file back for the next scan to take again — then press เริ่มใหม่ on the engine card, or
reinstall. Since v1.7.2 the engine carries a version block, manifest and icon like aetox.exe,
checksums.txt lists the hash of each exe on its own so a restored file can be checked against a
signed line, and that card names the likely cause instead of just the missing file.
Releases are signed: an ed25519 public key is compiled into the binary and the updater verifies
the signature over checksums.txt before it trusts a single hash. An empty or wrong key refuses
the update rather than falling back.
Not shipped as an app. The engine and the desktop package both compile and their suites run under
-race on Linux and macOS in CI; the browser pane is stubbed and packaging is not done. What
does ship for Linux since 1.6.0 is the engine alone — aetox-engine-linux-amd64 and -arm64 on
every release — as the half that runs on a remote host under ตั้งค่า › เครื่องระยะไกล, driven by
the Windows window over ssh.
1.0.0 is the Windows release. Until 2026-08-15 this line read "1.0.0 ships all three or it is not 1.0.0" — that criterion was changed by the owner, not met. Holding a stable Windows build behind a browser pane and an at-rest keystore that do not exist yet on the other two helps nobody already running it. Linux and macOS ship under the same bar, in a later release. See PLATFORM-SUPPORT.md for where the port actually stands.
go build -o desktop/build/bin/aetox-engine.exe ./cmd/aetox-engine # the engine, beside the window
go build ./cmd/aetox # the terminal console (runs the engine in-process when none sits beside it)
cd desktop
wails build # → desktop/build/bin/aetox.exe
wails build -nsis # with the installer
The window looks for aetox-engine.exe beside itself first, then falls back to go run ./cmd/aetox-engine inside a dev tree — wails-dev.bat builds it for you.
All of it happens on the same workbench you are watching. One window holds four rooms — slides, browser, files and terminal — and the Code door adds two more: Git, which lays out the uncommitted working tree with a per-file diff, and Timeline, the project's commit history a page at a time. The agent does not work behind a curtain and hand you a file at the end: it opens a room, works in that room, and you can reach in and change something yourself without waiting for the turn to finish.
It builds slide decks. Give it the subject and what you want out of it, and the deck comes back
as one self-contained .html file. Converting it is the app's work rather than the model's.
The deck is delivered when it opens in the slides room, not when it is exported — you page through
it and present it full screen from there, and the export bar is on that same screen. It exports as
.pdf, .png or .jpg: the PDF is the deck file itself through the renderer that draws it on
screen, so it looks exactly like what you were just looking at, and images come out one file per
slide into a folder of their own, named 01, 02, so a ten-slide deck sorts correctly everywhere.
A slide's box is 1280x720, which is 13.333 x 7.5in at 96dpi — exactly PowerPoint's widescreen
page.
Watch it work in a real browser. Not a headless scrape: a WebView2 window composited into the app, with an address bar, back/forward, DevTools, and eight device presets that resize the native window and zoom the page so CSS media queries genuinely fire.
The layer that drives it is ours. One read stamps a number on every interactive element on the page; click and type aim at that number. The agent hits the right control without a vision model and without guessing at coordinates, in the same tab you are watching.
Build websites and systems. Point it at a folder and the files room becomes a real place to
work: file tree, Monaco editor, unlimited real PTY terminal tabs, git, grep and glob over
the whole tree, plus diagnostics and symbol backed by language servers the app installs on
first use (gopls, typescript-language-server, svelteserver). Write a page, then open it and look
at the real thing in the browser room next door, without leaving the app to find somewhere to run
it.
Hand over a folder and get a file back. Point it at a directory of images, PDFs or recordings and ask for the thing you actually want. OCR (Thai and English), PDF text with the layout intact, and offline speech-to-text all feed the same conversation.
Read what the model cannot see. image_ocr runs Tesseract with Thai and English, so a
screenshot, a scan or a photographed form becomes text a 9B/35B model can reason about — no vision
model required, and the model that can see gets the image itself instead.
Delegate to a specialist. Pick @doc, @sheet, @deepresearch or @video off the +
menu — typing the characters does nothing, on purpose, since the day a pasted draft that merely
quoted @reviewer sent a whole brief to the wrong worker — and your sentence reaches that agent
word for word, not a paraphrase. The menu lists the team this chat hires from. Each agent is a
folder on disk with its own prompt, its own memory, optionally its own provider and model, and its
own private skills.
Give it a job, not a step. Work that takes twenty moves is planned before it is worked, and
todo_write puts that plan on screen while it runs, so what you watch is the order it chose
rather than a spinner. Up to four specialists run at once: task hands work out and returns
immediately, task collect picks it up, so three jobs cost the time of the slowest rather than
the sum. One that reaches a decision it should not make alone comes back as a question instead
of a guess. One still working when your answer arrives keeps working — you collect it by the same
id in a later turn, so the end of a turn is not a deadline. And before any of it runs there is a
planning stance that can read anything and change nothing, on an allow-list, so a tool added next
month is held back by default rather than slipping in.
Run on 2026-08-15, from one sentence — "find 20 CRMs a Thai SME could actually pick and give me
a spreadsheet comparing them": 6m 51s, two agents, 42 tool calls between them — 8 by the
assistant, 27 by deepresearch reading pricing pages, 7 by sheet — and one tool failure it worked
around. The handoff between the two was the baton this README describes: deepresearch left a
markdown report in the session's output folder and sheet was given the path, not the contents.
Twenty rows came back, fourteen with a real numeric price sorted low to high, and
six deliberately left blank with the reason written in beside them — quoted-only pricing, or
a page that would not state a figure. Every row carries the date the page was read and the link
the number came from, and says whether the page was opened in the browser or only searched. The
blanks are the part worth trusting: a table with no gaps in it is a table that guessed.
Ask it about your own past work. Every conversation and every tool run lives in local SQLite
with FTS5, so session_search across months of history is a query, not an inference — zero
tokens, Thai and English alike.
Have it build automations in n8n or Windmill. Connect an instance you host and the automation agent lists, reads, creates and updates workflows in it, and can start the server for you from a command you saved. Read the honest limit before you rely on this.
One switch on the wordmark moves between Assistant ("Use, remember, and create") and Code ("Build, debug, and ship"). It is the same binary, the same data directory, the same settings, memory and permissions — switching doors is not switching apps, and the app remembers which one you were in. The door also scopes the chat list in SQL, so a run of coding sessions cannot starve the other list.
| Assistant | Code | |
|---|---|---|
| Where it works | Your whole machine when no project is focused, or a project folder plus folders you add | The project folder you opened, plus folders you add |
| Rooms | Assistant · Capabilities · Projects · Specialist agents · Video work · Work | Code, with Git and Timeline as tabs |
| The right-hand panel | Available | Available |
The doors separate what the system carries, never what the AI is willing to do. The assistant has files and a shell and does software work with them; it does not hand a request back because it involves code.
A third door, Aetox Team, is built but not open in this build. It has one room, the Workroom — a run written down: the steps a job goes through, and which agent sits at each one. That page opens and says plainly that the work behind it is not built yet, rather than drawing a list it does not have.
The second door is a workshop, not a chat with a coding mode switched on. It opens onto one room, Code, rooted at the project folder you opened — in this door a project is a fence, where the assistant's project is a folder for conversations. Same binary, same settings, same permissions; what changes is the ceiling of tools and the room you are standing in.
What you get done in it, all of it on a workbench you are watching:
grep and
glob over all of it.The workbench beside the chat is where it happens — a real terminal, the browser the agent
drives, the file tree and a single-file editor, git status against HEAD with +N −M, a map of the
code, and a pull-request room. Behind the file tabs a Monaco editor; behind the terminal tabs a
real ConPTY.
What is deliberately not on this desk: no document or spreadsheet writer, no OCR, no PDF or audio reader. A deck about code is the assistant's door.
Seven agents ship — doc, sheet, github, automation, deepresearch, editor, video —
and hiring an eighth is dropping a folder into <DataRoot>/agents/. No release, no plugin API, no
restart.
An agent's folder is its whole identity: AGENT.md (who it is, what desk it sits at, which tools
it may narrow itself to, which model it pins), MEMORY.md (what it has learned), STARTERS.md
(how an empty chat with it opens, per language), and a private skills/ folder no other agent
can see.
That folder is also where the difference between a clever assistant and a company sits. Each agent pins its own model, at its own provider, so the one that opens twenty pricing pages can run on something cheap while the one that has to weigh what it found runs on something strong, and the bill follows the work instead of following the hardest task in it. Each keeps its own memory, so what the deepresearch agent learned about a source does not leak into the document agent's judgement about a contract. A single generalist has one model, one memory and one set of tools for every job it will ever be handed — and no way for you to add an eighth colleague to it.
You can delegate to one — the assistant calls task and up to four run concurrently — or you
can talk to one directly, in a session bound to its tools, its memory and its prompt. @name
from the composer is the third door: your sentence arrives verbatim, mention included, because a
paraphrase is where the request goes wrong.
Agents are hired through teams. A team is a list of agents bound to one desk
(<DataRoot>/teams/<name>/TEAM.md), each side of the app — ผู้ช่วย and โค้ด — has its own teams,
and a chat hires from one team for its whole life, or from none. The app seeds one team,
ผู้ช่วยในคอมพิวเตอร์, and after that it is an ordinary file you can rename, trim or delete. Teams
are managed under ตั้งค่า › ทีม; the chip beside the composer says who answers and which team
it hires from.
Agents never call each other. The star has one centre; multi-step work is a conveyor through the
assistant, and the baton is a file path rather than the content. Separately, two sub-agents
(explore, general) are internal helpers — a fixed set, not extensible,
deliberately, though since 1.6.0 you may tune one: its model and provider, its prompt, its step
ceiling and its look, never what it can reach.
Aetox remembers across sessions, and nothing is written without you approving it.
The memory tool cannot write. It queues a proposal. Separately, a summarizer reads the tool-run
log with no model call at all, clusters repeated failures by agent + tool + normalised error, and
proposes a lesson once the same mistake has happened three times.
Everything lands in one review queue in Settings, and each card shows the body, the agent's stated
reason, whose memory it would go into, and — for a replacement — the line it would overwrite.
Approve or discard. What is kept is plain markdown you can open, edit line by line, or forget in
place, and every decision is recorded permanently, so "why does it think that?" always has an
answer. One switch turns the whole thing off. Since 1.6.0 the door also remembers what you decided:
a fact you declined is not proposed again in other words, one already kept is not asked twice, and
the model is told what is pending and what was refused. Memory is kept per desk — MEMORY.md is
the assistant's, the Code desk has its own file — and every label says which desk reads it.
It takes effect from the next session, not this one — a mid-conversation prompt change would invalidate the provider's prefix cache, which is the same reason the tool block never moves.
The other half is standing instructions: your own always-on markdown files that ride into every desk, every project and every agent. What you wrote, and what it worked out, are kept apart on purpose.
A desk is the tool ceiling of a session. Three ship — assistant, coding, specialized —
and a session's desk is fixed for its life. Desks are also what MCP servers and external
connections are placed on, which is how a tool installed for one kind of work stays out of the
others.
Two processes. Since 1.6.0 the window is a screen and aetox-engine is where everything
happens — the model loop, the files, the shell, MCP, the database — spoken to over one socket even
on the same machine. The window keeps the things that are the machine's: the browser pane, computer
use, dialogs, the voice, and every provider key. The engine never holds a key; it hands each request
to the window to sign, and the window signs only for hosts it knows belong to that provider. That
split is what lets the engine run on a Linux host over ssh with the same code and a longer wire.
Where it may go. With a project focused, the workspace is that folder plus any folder you add
— added folders get read and write with no prompt, the same rights as the root, because a second
quieter tier would be a rule you never agreed to. With no project focused, the workspace is the
machine, and writes land under output/<session>. One function resolves every path, symlinks and
all, and there is deliberately no second check anywhere else.
Approval. Three levels, and one gate every tool call goes through — built-in tools, shell and MCP alike.
| Level | What it asks about |
|---|---|
| Ask | Anything that is not a plain read inside the workspace |
| Unsafe only | Deletes, git, shell, and anything touching a path outside the workspace |
| Full access | Nothing. There is no carve-out. |
MCP tools confirm at every level regardless, because their behaviour is defined by somebody else's server.
Shell commands are path-contained, not pattern-matched: a path hidden in a quoted argument,
behind a flag, behind a redirect, or behind %VAR% / $VAR / ~ is still resolved and checked,
and a command the scanner cannot read — $(...), backticks, -EncodedCommand, FromBase64String
— is refused rather than guessed at. Every command run is appended to a 0600 audit log.
Refused to every file tool, in every mode: .ssh .aws .gnupg .azure .kube .netrc
.git-credentials .config/gh .aetox, the Windows Credentials and Protect stores, Chrome /
Edge / Firefox / Brave profiles, and Aetox's own credentials.json, oauth.json,
account.json, mcp-servers.json, screen.json (the remote hosts and the token that admits the
window to an engine) and browser profile. Folder-picking refuses them too, so it fails at the door
rather than as a confusing tool error later.
Your data.
| Where it stands | |
|---|---|
| Chat history, tool runs, produced files | On your disk, in local SQLite and plain folders |
| Browser data (history, cookies, session) | Stays on your machine only — no server of ours sits in between |
| Cutting the cloud off entirely | Your data stays on your machine and in your country — run through LM Studio or Ollama and not a single byte leaves |
| API keys | Their own file, 0600, DPAPI-wrapped against your Windows account. Off Windows there is no encryption at rest — that is stated rather than implied |
| Secrets in logs | Stripped through one registry into all three sinks: debug log, shell audit log, and the buffer the bug-report form reads |
| MCP secrets | ${env:VAR} indirection, so a key never lands in the settings file |
| Taking it with you | Export any chat to .md or .json, and import a .json back into any Aetox |
| Bug reports | The app transmits nothing. It prefills a GitHub issue, already scrubbed, and you read every line before sending it from your own account |
There is no server of ours in the middle and no analytics. Using a cloud provider means that provider sees what its API normally sees, and nothing is routed through us.
A tool count is not a reason to use anything, which is why this is down here.
35 tools reach the model on a fresh install — 34 from the engine and browser, which the
window lends across the wire (§248); computer joins only once you switch it on. A default
assistant session carries fewer, because a desk narrows the set. They cost about 10,700 tokens on
every request before you have typed anything — the engine's 34 are about 9,900, against a ceiling
of 10,400 tokens and 48 tools that a test enforces on that block, and the browser's definition is
another ~830. Twelve of them are packed — one name in the block, several verbs behind it —
which is why the list got shorter in v1.5.15 without anything being taken away. Re-measured
2026-09-17 on v1.7.2.
| Group | Tools |
|---|---|
| Files | change (write · edit · append · batch · delete) read search (list · glob · grep) |
| Running commands | computer (list_apps · read · capture · focus · click · type · close — only once switched on in ตั้งค่า > การใช้คอมพิวเตอร์) desk_terminal git shell (run · output · kill · list) |
| Handing back files | asset_find doc_write sheet_write video (new · check · render) |
| Reading media | image_make media_read (image · video · audio) pdf_read video_project |
| Web | browser (open · read · click · type · wait · back · scroll · capture · tabs · dialog · console · network · hover · drag · key · upload) media_fetch web_fetch web_search |
| Code work | codebase (errors · symbol · impact · map · trace · design) rename |
| GitHub | github (search · repo_summary · list_files · read_file) pr (list · read · checks · create · comment) |
| How the assistant works | ask_user calc desk (open · list · close · focus) memory plan (write · amend · read · step · report) plugin_install session_search skill_view task (start · collect · answer · message · plan) time todo_write |
That table is generated from the registry the model is actually handed
(go test ./internal/engine -run TestPrintReadmeToolTable -v, plus the two the window lends),
because a hand-kept list of what a program contains is a second source of truth for a question the
program can answer — and this one drifted for months, still naming tools that had been folded into
shell and github.
Connecting an automation engine adds one more packed tool — n8n (list · read · create · update ·
activate) or windmill (workspaces · list · read · create · update) — and nothing until then:
a tool with no account behind it is withheld rather than shown and refused.
Growth goes where it costs nothing. A skill is a markdown document, not a tool: the prompt's skill index
carries the names, grouped by area, and skill_view returns one body, so installing three hundred
leaves the tool block exactly the same size. MCP servers are placed per desk and per agent, so a server added
for video work is absent from an ordinary conversation — not hidden from the model, absent. Office
writers reach only the specialized desk, so the assistant delegates for a .pptx rather than
carrying three tools it rarely needs.
27 providers, and the window shows every one — OpenAI · OpenAI-compatible (your own endpoint) ·
Anthropic · Gemini · DeepSeek · Qwen · Z.ai · OpenRouter · Codex · Groq · Mistral · Kimi ·
MiniMax · Xiaomi MiMo · Xiaomi MiMo Token Plan · xAI · Meta · ThaiLLM · ModelScope · NVIDIA · GitHub Copilot · Kilo · Ollama Cloud ·
OpenCode Zen · OpenCode Go · LM Studio · Ollama · and the built-in aetox. ChatGPT (Codex), GitHub Copilot and OpenRouter
sign in; the rest take an API key or a local server address. The catalogue and the picker used to disagree; they no longer
do, because a provider the engine knows and the window hides is one nobody can reach.
Local models are treated as first-class: Aetox asks LM Studio and Ollama which model is loaded rather than which exist, streams the answer and the reasoning, really calls tools, and counts tokens into the same statistics. You can switch provider or model mid-conversation and the full context follows — tool calls, tool results and compaction summaries, not just the visible text. One provider is active at a time; Aetox never silently reroutes your turn to a different paid one.
Connect an n8n or Windmill instance you host and the automation agent can list, read, create, update, and — n8n only — activate workflows, and start your server from a command you saved.
It cannot run a workflow and see the result. There is no execution API call anywhere in this codebase; the closest thing is the agent clicking Execute in the vendor's own editor through the browser tool, which is not a verified run. Windmill has no activate either, so a flow it creates is saved and inert until you trigger it yourself. The agent says so out loud rather than implying otherwise, and a test exists whose only job is to keep it saying so.
There is no scheduler, and there will not be one. Aetox has no cloud, so a schedule would silently depend on your laptop never closing. n8n and Windmill are the clock; Aetox is the hands.
An answer cut off by the output-token limit is continued. A reply that hits the ceiling used to reach you stopped mid-word, with nothing anywhere asking for the rest. The turn now carries on up to three times and appends to what is already on screen, so one answer is watched being written rather than vanishing and starting over.
A tool call that names the same argument twice is refused. A model asking for three searches
sometimes writes two of them into one object — {"query":"A","query":"B"}. That is valid JSON, so a
parser keeps the last and drops the first without a word: one search never runs, and the answer
reports on it anyway. Aetox rejects the call and tells the model what was actually wrong with it,
rather than letting a silent loss reach the answer.
A provider that returns nothing is an error, not an empty answer. A turn 350 seconds in with eighteen tool results behind it once died on a round that came back without a single frame of text. That round is replayed twice — with whatever streamed taken back first — and only then does the turn change the question instead of asking a fourth time. Everything the turn had already done stays in context either way.
What a model can do only ever narrows the toolset, never widens it. A model the catalogue has never described keeps every tool; one the catalogue says cannot call tools is narrowed. Wrongly withholding tools turns an agent into a chat window, so doubt is resolved in one direction only.
The rules are in BENCHMARK.md, and its one standing rule is that a number which has not passed them may not appear here or on the website.
The dangerous number is the flattering one, because nobody audits a figure that makes them look good.
Aetox. The two size rows and the two test counts were re-measured 2026-09-17 on v1.7.2; assembling a turn is from 2026-08-13, and the ⁽ᵈ⁾ rows from 2026-07-27 on v0.9.2 — before the engine became a process of its own, so the process count in particular is one short of today.
| What you download | 34.6 MB installer |
| What ends up on disk | 83.4 MB, two files — aetox.exe 50.9 MB + aetox-engine.exe 32.5 MB |
| Assembling a turn | 0.32 ms · 174.9 KB allocated |
| Go tests | 3,648 across 62 packages, 0 failures |
| Frontend tests | 2,106 across 205 files, 0 failures |
| First launch (cold) | 1.77 s ⁽ᵈ⁾ |
| Every launch after | 0.53 s ⁽ᵈ⁾ |
| RAM committed | 252 MB ⁽ᵈ⁾ |
| Processes | 7 ⁽ᵈ⁾ |
⁽ᵈ⁾ Measured 2026-07-27 on v0.9.2 under the rules, and not re-measured since. They are dated figures rather than current ones, and they are here rather than deleted because they did pass the rules on the day — which is the whole difference between an old number and a bad one.
Two things that number honestly. Assembling a turn was 0.12 ms and 96.2 KB when the block held 27 tools; it is 0.32 ms and 174.9 KB now that it holds more. That is a real regression, and it is still three ten-thousandths of a second — the time you wait is the model thinking. And the disk figure went from 48.5 MB in one file to 83.4 MB in two: the engine is now its own executable, and the window still links the engine package for its types and forwarders, so the split added a binary without yet shrinking the first one. That is a real cost of §248 and it is written down as one. The Go suite is green on Windows; CI on Linux and macOS is red — 16 tests on 2026-09-13, the port's unfinished edges (shell path rules, WSL, the window tools) rather than the engine. Since 2026-08-15 those two jobs are reported rather than gating — Windows is what ships, and one shared verdict meant every Windows push went red for a port's unfinished edges until nobody read the colour at all. The failures are still on the run page and still to be fixed; what changed is that they no longer hide the platform that is done.
Against Zed, the harder ruler — native Rust, with a reputation for being light.
| Aetox | Zed | |
|---|---|---|
| First launch (cold) | 1.77 s | 2.12 s |
| Every launch after | 0.53 s | 0.53 s |
| RAM committed | 252 MB | 471 MB |
| Disk | 83.4 MB | 419 MB |
Both columns except Aetox's disk figure were measured 2026-07-27 on the same machine under the same rules, and neither has been re-measured — Zed is no longer installed here. A tie on warm launch with a native Rust editor is the result worth having; treat the row as dated rather than current.
The rest of this category ships 240 MB to 1 GB because an Electron app brings its own copy of Chromium. Aetox uses the WebView2 that Windows already has — and being straight about it, WebView2 is Chromium, so the memory it holds is not a win over Electron. The win is that you are not handed a second browser to store.
Disk — download the portable zip,
unpack it, and add up the two files inside: aetox.exe 53,339,136 bytes and aetox-engine.exe
34,120,192 bytes, 87,459,328 together. Anyone can reproduce it in a minute. It replaces the 48.5 MB
single-file figure measured on 2026-08-25 on v1.5.7, which was correct then and is not now. Competitor sizes are measured after install from the install folder, never
taken from a download page, and never from a folder holding user profiles or caches.
Launch, RAM and process count — bench.ps1 -Start, empty project, median of 5 runs after
discarding the first, read after 60 seconds settled. A true cold launch needs a reboot first,
because Windows keeps the app's files in its file cache afterwards.
Assembling a turn — bench.ps1 -Engine, median of 3 rounds.
What was removed from this section. An earlier version of this README published "97% of input tokens came from cache over six consecutive messages" and local first-token times of 1.42 s and 1.75 s. Neither has a source in this repository — no test, no log, no BENCHMARK entry — and the machine those local numbers describe did not have LM Studio installed. They are gone rather than date-stamped, because the rule above does not have an exception for numbers we would like to keep.
A skill is only worth shipping if the same model does the same job better with it. So the
question was measured, not argued: the same model (gpt-6-luna, gpt-5.6-terra), the same 39
tasks, Aetox with its skills on and with every skill placed off, and Codex CLI and OpenCode on the
same model beside it. Every task starts from a fresh folder and a fresh data folder, is scored by a
program rather than a person, and was proven scorable by a reference solution first; the hidden
tests never enter the folder the model works in. Task sources: SlopCodeBench, SWE-bench
Multilingual, Web-Bench, CWEval, SlidesBench, Terminal-Bench (adapted), and our own where no public
set covers the job (Thai tax invoices, PromptPay, an outage with only the logs).
Only differences that held on every repeat are quoted here:
| Task (runs) | Skill | Skills on | Skills off |
|---|---|---|---|
| Thai slide deck, Luna (3) | aetox-slides | 98 | 49 |
| Thai slide deck, Terra (2) | aetox-slides | 97 | 29 |
| Thai tax invoice, Luna (3) | aetox-th-locale | 100 | 94 |
| Thai tax invoice, Terra (2) | aetox-th-locale | 100 | 85 |
Save a /command preset (6) | aetox | 5 of 6 | 0 of 6 |
And what did not move, stated with the same weight: a design-token page (Terra), a leaked-token cleanup, closing a branch and three unclear functions scored the same with skills on and off; a safe database migration scored slightly lower with the skill (Luna 83 vs 92, Terra 88 vs 100). Across all 35 scorable tasks, run once each on Luna, skills on scored 81 and off 77. Against the other harnesses, on 19 shared tasks run once: on Luna Aetox 86, OpenCode 82, Codex 75 — but Aetox with skills off also scored 86, so that lead is the harness, not the skills; on Terra Aetox 85, Codex 87, OpenCode 91, a loss we have not closed. Every task, scorer, run count, excluded run and changed threshold is in SKILL-BENCH.md.
The core is in place. Release notes · roadmap.
Three things it does today that are worth knowing about:
.html file makes
Aetox load it at desktop and phone width, light and dark, and what needs fixing comes back with
the write: script errors, text under AA contrast, a page wider than the screen, a dark view that
stayed light, Thai whose marks float or overlap the next line, controls too small to hit. The
assistant no longer spends rounds opening the page to look.Next — one provider chain that switches accounts when a plan window is spent · external agent programs as engines (§258) · the personal secretary on the user's own Apps Script bridge · a same-state RAM round for every app in the tour's last scene.
Measurements and test results live in docs/reports/ — Benchmark rules ·
Test report by module · Token audit ·
Where every published number lives ·
Research & reasoning evaluation ·
Harness database pilot ·
What the skills are worth.
Using and shipping: First-run tour · How a release is cut ·
Platform support · Roadmap.
Architecture, decision records and design standards are kept off this repository (docs/internal/, not published).
There is a Facebook group for questions, ideas, and the kind of half-formed problem that does not fit in an issue yet: the Aetox group. Bugs are still better filed as issues, because an issue carries the version and the log with it.
Open an issue. The app has a door for this:
Settings prefills a GitHub issue with your version, install channel and OS, folds the recent
internal log into a <details> block with secrets already stripped, and hands it to you to read
before you send it from your own account. Nothing is transmitted by the app.
Aetox is written by one person. It exists because a model that can only produce text is half a tool, and the missing half — hands, permission, and a place to put the result — is an application problem rather than a model problem.
Aetox is free to use and is not open source. From v1.3.0 it is under a proprietary licence: install it on as many machines as you like, use it for commercial work, read the whole source and audit it — but do not modify it, redistribute it, rebrand it, or sell it. The source is published to be read, not to be built on.
Your own extensions are yours. Skills, agents, prompts, configuration and MCP servers you write are your property, and selling them is expressly permitted (LICENSE §4). That is what the extension points are for.
The name "Aetox" and the logo are trademarks and are not licensed to anyone else. Third-party components keep their own terms and are listed in THIRD-PARTY-NOTICES.md — none of them is GPL or AGPL.
Earlier releases keep the licence they shipped under, permanently: v0.7.1 and earlier are MIT, v0.8.0 through v1.2.4 are Apache-2.0.
Aetox was not born to compete with anyone. It exists to stand where the market has a gap — not to be one more agent framework, and not to lock anyone into anything.
📧 phrmsawanachyphl@gmail.com · ❤️ Support the project
© 2026 Chayaphon Phromsawana · All rights reserved · Licence · Third-party notices
Replace {MCP_ENDPOINT_URL} with this MCP’s endpoint URL (from its repo or docs above). No API key — you connect directly.
Tool
OS
Config file: ~/.cursor/mcp.json
{
"mcpServers": {
"mcp-server": {
"url": "{MCP_ENDPOINT_URL}"
}
}
}Paste into mcpServers in the config file. Restart Cursor after saving.
If this MCP is also published on mcpchannel.ai, you can subscribe from Browse and use the gateway config there instead.