back
90 comments
Linking to a github repo for a binary release (no source code related to the agent that I could see) is a bit iffy IMO. You should clarify your intentions or link to something else. Might confuse folks.
if you're interested in an open source agent that uses a minimal amount of CPU and RAM and has source that is easy to audit:

https://github.com/smol-env/smol

here are traces from an agentic task around using duckduckdb

comparing CPU and RAM usage of the whole container over time w/ OpenCode, hermes, pi, codex, smol

https://x.com/__tosh/status/2086882367126286466

https://x.com/__tosh/status/2086882204060160350

smol is very minimal only using stdlib (in this case it is the go version but you can also take a look at implementations in python, clojure, php)

also no need to trust these bench runs, you can just run your own

(any OpenAI Responses API compatible endpoint works, if your endpoint does not support 'custom' tools you can have your agent change the smol implementation to use 'function' tool implementation instead)

that said: be aware that smol does not come with any system prompt and does not load agents.md files by default

some older not so strong models benefit from a system prompt and guidance in agents.md that complements them

that said 2: system prompt and or loading agents.md automatically is easy to add though if you want it

I will leave this here: https://usehax.dev/ GitHub repo: https://github.com/OleksandrChekhovskyi/hax

This is a coding agent implementation I am working on, which delivers what this promises (at least on the "lean" part), except it's actually fully open source, and even more lean (few MBs of runtime memory usage).

MIT-licensed, written in C, multi-provider / multi-model, minimalist approach to system prompt and tools (think kinda like pi, but with a bit more "batteries included", like subagents and background tasks out of the box), polished presentation, inspectable (usable transcript view), etc.

nice, good to see more contributor in this space
I was kind of excited for this until the binary blob. You want me to give your agent binary god access to my computer, and I am not even permitted to see the source code or use my own supply chain security hardened rust compiler stack? What a joke. Hard pass.
we put it in the repo README, will add migrate more into public repo as soon as possible.
It has to be 100% open source public code or we will have no way of proving this is not malware, or secretly swapped out for malware later when your CI/CD system or laptop is compromised. Supply chain attacks happen all the time and with closed code no one will be equipped to spot it when it happens.

Also, security aside, engineers want the freedom to modify and experiment with the tools we rely on.

Tools like this are too important to be closed. Do you want to be Internet Explorer or Firefox?

Linking to a binary is iffy from a security perspective. Linking to a GitHub repository is exactly what HN should do.
> “while taking the time to figure out how open source should work in the agentic era”

I can’t even guess what this means

generally the challenge now is that 1. how to deal with PR spams by AI bots 2. how to make the project sustainable especially when one has no distribution. when anyone can insta remix and re-package and re-sell your hard work. For knowledge sharing open source it is ok, but if you are serious about what you built, this is question needs to be answered first before make it a true community effort.
I have bad news for you. No matter what you do here it is probably not sustainable. Competing commercially with people who have so much more resources and ability to build is not sustainable. If it's open source you need adoption, and shouldn't care if people want to steal and remix your work as that will drive more adoption.

But from a fellow engineer's perspective we don't need another commercial solution in this space. Every major tech vendor is working on harnesses, and we will get ones that will run circles around yours for free. The open source ones that get popular will have mit/freebsd/etc permissive licenses. (or no one will adopt them) But coming at it from the angle you are indicates you haven't reasoned well about this. Please don't hurt yourself and those around you by starting a startup on this idea.

there are still many bad players in the industry, i wouldn't mind sharing it with trusted group. But I am not yet strong enough with deal and handle all those yet.
I understand that claude-code takes a lot of memory and that's bad. However, harneses are simple loops, in theory should take very little memory even if written in python or typescript. See for e.g. pi agent
I love pi and share many vision and value with it. But my view on harness is that it is to capturing the structural mechanism with llm interacting the world. They are a dynamic duo evolving together. The technical depth will continue to grow (e.g. /goal being the new primitive, multi-agent collaboration is basic need)

Yes the core part is a simple loop, but we have all built toy compilers, inference engine, browsers (it is just a curl command eth) etc. The core algorithm is supposed to be simple.

Had a discussion recently: https://x.com/NoCommas/status/2086568454434537710?s=20

no source code
> We care about the harness, not the model or the prompts.

I wonder if this is a viable approach; after all frontier model providers are betting on the opposite.

Then again, they bundle their harness and offer subsidiary pricing - so maybe they themselves aren’t sure if models are as important.

my view on harness is that it is to capturing the structural mechanism with llm interacting the world. They are a dynamic duo evolving together. The technical depth will continue to grow (e.g. /goal being the new primitive, multi-agent collaboration is basic need) So it is here to stay. And it is just our focus as we don't have enough resource (yet) to improve the model and I think prompts belong to the user.

Had a discussion recently: https://x.com/NoCommas/status/2086568454434537710?s=20

How good is it to work on building games, compared to existing agents? I am building my own game?
They are a bit weird with game development at the moment.

They can one shot entire games, with relatively minor issues.

And obviously asking for small code snippets and integrating them yourself has been well supported for five years.

But in Agent mode... not so much. I was asking frontier models to make simple changes to my Pong game (you know like the one from 1972) and it constantly failed to make simple changes or would break something else in the process.

The main issue is that they can't see what they're doing. Actually one of the agents tried playing the pong game by screenshotting every frame, and it ran for about 20 minutes before I realized what it was doing, and told it to calm down.

It takes about 10 seconds to process an image, so it was running the game at 0.1 frames per second... 600x slower than realtime. The technology is not quite there yet.

If your game is something turn-based though, with discrete States and well-defined transitions between them, they can help out a lot more with that.

You aren't building your own game if you have a chatbot do it for you.
we have a detailed launch thread explaining and show case exactly this! https://x.com/NoCommas/status/2086835536598351955
"Is there telemetry? Yes, and it is opt-out: set ANTE_TELEMETRY=off"

Opt-out every time is unacceptable. If you opt out once it should be enough. One accidental execution path without the right environment and you're spewing telemetry to spaghetti knows where.

> One accidental execution path without the right environment and you're spewing telemetry to spaghetti knows where.

That would also be true with opt-in by environment variable, though.

called out the most asked questions - where is the source - telemetry opt-in/opt-out in README of https://github.com/AntigmaLabs/ante
Considering that ripgrep, git, and, you know, other dev tools are part of the toolbox, then why ship them inside this executable? And, furthermore, if you ship them, then why stop there?
so it is more for being self contained and works out of box if being deployed in a bare linux environment. we tried shell out to `rg` it didn't work very well and instead spending time handling the args parsing and jugging string output, we decided to spend time on building Grep natively for agent. It is just a start~

as for why stop here yep, the goal is to be able to find the best sweet spot in being self contained v.s. all-in-one bloat ware. for example, we still use a bundled `tmux` skill for the orchestration.

so it is just llamafile with few additional application in bundle?
it is more like an actual harness with a managed and pinned llama the control is inverse.

Harness is compute and Model is data

Wow. Great concept!
> Is there telemetry? Yes, and it is opt-out

Fuck you. Will there ever be a decent agent where the answer to this is "No"?

this would go very hard with a lightweight gui
yes, the goal is to perfect the `ante serve` so it is easy to build gui. We are building one internally to test the protocol version
Hi HN, I'm Mohan from Antigma Labs. Ante is a coding agent that ships as one self-contained ~15MB binary: the TUI, an embedded ripgrep, local PDF/OCR, and a natively managed llama.cpp engine are all inside. No runtime dependencies, no node_modules, no account.

- Ante installs a pinned, checksum-verified official llama.cpp build matched to your machine (Metal on Apple silicon; CUDA, Vulkan, or CPU on Linux) and handles upgrades when the pin changes. - It discovers GGUF files already on disk (~/.ante/models, the llama.cpp and Hugging Face caches), attaches to llama servers already running on local ports, and estimates RAM/VRAM from model size and context window before anything loads. - `ante --offline-model /path/to/model.gguf "prompt"` boots the server, runs the session, and shuts it down. `/offline-mode` does the same interactively; `ante serve --offline-model` loads a model once for many clients. - No API key, no account. Once the model is on disk, inference needs no network at all; set ANTE_TELEMETRY=off and no telemetry is exported either.

On capability, we'd rather publish the number than oversell: we benchmark local models with the same harness and auditable runs as frontier ones, and Qwen3.6 27B (a 17 GB download) scores 56.2% on Terminal-Bench 2.1 across 445 trials (live results: https://antigma.ai/eval). That's a real gap from frontier models. The design bet is that you mix: hosted providers and local live in the same catalog, `/providers` switches mid-session, so sensitive repos or high-volume work go local and hard problems go frontier.

Hosted models work with your own keys or subscription. But nothing about trying Ante requires signing up for anything: download the binary, point it at a GGUF.

Offline mode is under active development and has rough edges with the overview at https://ante.run/local/overview. I'll be in the comments.

Here you say:

  > Ante installs a pinned, checksum-verified official llama.cpp
But in README:

  > Ante ships its own inference engine
May I suggest you use the first phrasing in both places. I took it as Ante devs had written their own engine and I doubt I'm the only one.
Not sure why this was dead but I vouched. It would be nice if telemetry was opt-in, otherwise this looks awesome and can't wait to try it!
Where is the source code?
Why not make it Actually Portable Executable using Cosmopolitan Libc, like Llamafile, to make it run on Windows/Linux/MacOS? Why don't you support Windows with CUDA?
Many people have slow computers, but for agent it is no problem. Only run LLM slow too.
Opt-out telemetry is a hard no for me, sorry.