back
164 comments
Interesting to see they shipped an "anti-slop" taste skill:

> description: Anti-slop frontend skill for landing pages, portfolios, and redesigns. The agent reads the brief, infers the right design direction, and ships interfaces that do not look templated. Real design systems when applicable, audit-first on redesigns, strict pre-flight check.

> - *PREMIUM-CONSUMER PALETTE BAN (mandatory, second-most-recurring AI-tell):* - For premium-consumer briefs (cookware, wellness, artisan, luxury, heritage craft, DTC home goods, etc.)... - Backgrounds: `#f5f1ea`, `#f7f5f1`...

> Landing pages and portfolios are *visual products*. Text-only pages with fake-screenshot divs are slop.

https://github.com/yc-software/qm/blob/7f2c916360f1797a8ff2a...

> Em-dash (—) is COMPLETELY banned. It is the LLM's signature stylistic crutch and it is the #1 visual Tell in production tests. There is no "limited use" allowance, no "natural language frequency" allowance, no "in body copy is fine" allowance. None.

I guess the em dash is really dead.

All due respect to the YC folks, but having a skill with 22,069 tokens is a major skill issue. Yes, ironic.

My slop control skill is a thousand tokens. Biggest problem I see with the skill is that everything is prompted via negativa. smh. sorry to be judgmental but it's hard to trust a harness that comes shipped with a skill like this.

Doesn’t this just lead to a new “basin of tastelessness” that, sure, looks different from current slop, but is itself just eventually slop all the same?
According to the license, the source of it is actually this https://www.tasteskill.dev/

Which, in my opinion, feels like slop.

They all use the same pill glowy status light header eyebrow to immediately tell you they have no actual design skills and follow the herd.

Posers took over the industry.

lol this is hilarious - the “make no mistakes” of design

You can’t prompt an agent to have taste.

Love seeing this direction along with Buzz.

The hardest problem in multiplayer agents, at least for us, has not been the agent loop. It is scoping and QM's per-person scopes plus shared rooms is a sane answer for a company-wide assistant.

I build in the adjacent lane, AQ (aq.dev), a multiplayer coding harness where teams run Claude Code and Codex together), so seeing YC ship "a multiplayer agent harness for work" is validating and a little surreal.

Is Hermes the best openclaw like agent as they mention running it before?

Also, what are power uses really using openclaw like systems for?

Still figuring it out, but it's been really convenient to have an always-on agent that has access to internal systems and can be triggered by webhooks. Some examples of what we use it for:

- automatically fixing simple CI failures

- getting production alerts and automatically creating RCAs and a fix PR

- periodically checking slow DB queries and finding ways to speed them up.

- creating charts to answer one-off questions about our data

I've tried using it as an on-the-go coding agent as well, but found I prefer more interactive agents, so I can see what the code looks like.

Hermes is huge and packed with features you probably don't need. I prefer smaller one I can extend as necessary, there are so many on github now and it is fun to test them but have been impressed with dirge (https://github.com/dirge-code/dirge) not affiliated.

I have one reading my second tier RSS feeds and newsletters and giving me news/market updates filtered for things important to me

Hermes is what I was using but I still found it annoying I often wanted to operate 1-2 levels deeper.

Yesterday I just decided to try writing my own version (100% just for me, not open source, no monetization plan, incredibly custom) and I've been enjoying working on it so far (I know, I know, it's been a day, honeymoon period and all that).

Part of it is I like building software (even if I'm not writing every line) and part is I like having full control. Turns out (and many people have said this) the basic agent loop really isn't all that special. There are a million levers to pull and what not outside the base loop but that the fun part for me. Trying out different ways to add on to the core concept.

I'm really enjoying being untethered for things like "how will I monetize?" or "how do I make this generic so others can use it?". If I need functionality I just add it in, I don't need to make it infinitely pluggable, etc.

All that said, I'm thankful to things like nanoclaw and then Hermes for exposing me to the core ideas. I just want to put my own spin on it.

Can someone give me an example where this is truly and uniquely useful? I've seen so many of these things and I can't tell if part of it is like mostly for enabling nontechnical folks to do more things or if it's some unlock and additive value. At end of day beneath it all things are just prompts, and then you can provide tools and context, but even then that's not even something I've found that useful to keep because context windows are still limited and bias propagations often need to be constantly corrected, especially when digital artifacts are not perfect representations of the real world.
We’ve been deploying QM into customer AWS accounts (official Terraform path). Happy to answer AWS/ops questions. DIY vs managed cost sketch: https://digitize.llc/qm/calculator/ · writeup: https://digitize.llc/qm/
I gave an agent its own Slack channel and it started scheduling meetings with other agents without me. I've never felt more like middle management
It's fascinating to see new UI primitives and concepts get invented in the LLM era. The sea of creativity makes it hard to even understand most of what each new app does, and nobody describes them well. When I went to the Hermes agent web page, I was left with zero clue about what it did or what it could do. It took a bit of digging to find the right part of the qm page that helped me grok what was going on.

I've become attached to Orca (yc-backed) for managing coding sessions in the past week, but some sort of postgres session db is what's really lacking. So, maybe it's time to try qm.

Aren't there a ton of products already doing this? Why not just use claude Cowork? Surely they're simpler/better/more featureful/developed than the alternatives here? What advantage does this have? Would love to see a 'QM vs Cowork' comparison!
How is QM trying to differentiate from HarnessRouter.ai, which launched on Launch YC about 10 days ago? https://www.ycombinator.com/launches/RpL-harnessrouter-bring...

Both seem to provide a unified interface across different harnesses/models. HarnessRouter exposes that as a developer API for embedding agent-powered features into products, with Task / Run / Session / Streaming / File / Artifact / Renderer contracts, plus tracing around harness/model behavior over time.

Is the main distinction that QM is a company workspace, while HarnessRouter is an embeddable harness layer for product teams and enterprises?

This was literally in YC's Request For Startups for Fall 2026: https://www.ycombinator.com/rfs#multiplayer-ai
What is yc software?
I love the concept to use different harness frameworks in qm but a true multiplayer harness needs to support other agents and any MCP clients, including Cowork.

Making agents multiplayer is mostly a context problem. You could be using ChatGPT or a Slack bot, or a web interface and the agent needs to know you, your conversations in Cowork etc. so it can enable multi channel collaboration with your agents and your colleagues. We're working on it at https://lobu.ai

> We take contributions as human-written text, not code — see CONTRIBUTING.md. Describe the change you'd like informally in a .txt or .md file in adrs/, and if we're aligned we'll handle the implementation.

Interesting approach to open source contributions. Closer to feature requests at that point?

So that explains why there is so much low effort cold outreach on LinkedIn from YC founders these days.

It's getting ridiculous the amount of unsupervised agents doing active inbox management on things that should be personal relationship work. I'm so tired of it.

I'll need to explore how they're doing org wide context and security for sure. This seems extremely complementary to my own coding tool which currently gives the best AI interface for individuals.

Would be cool if in a few epic tickets i'll have both their org wide architecture AND a productive individual coding interface :)

In a tangential earlier release, Gary Tan's own gstack:

https://github.com/garrytan/gstack

Use Garry Tan's exact Claude Code setup: 23 opinionated tools that serve as CEO, Designer, Eng Manager, Release Manager, Doc Engineer, and QA

If I had applied to yc, I would have thought they had "distilled" my startup idea xD
Multiplayer agent harness is useful only when the context and AI output does not overwhelm the team space and works asynchronously and background. If you think about it, that could be just boring job scheduler, isn't it?
A video would've been nice to walk through how it's actually used. I don't mean just a walkthrough after a fresh install, but an instance that's actively used.
looks like an internal tool that yc rushed out the door to minimize any most lost ground to Buzz. That said, I'd be curious what folks think comparing these two tools.
Pretty cool - I’ve also been doing something along the lines of this with lumifyhub but it includes docs and boards natively
it says "Each deployment runs in the operator's own cloud account"

But feels like its written to run on one mac/vm and carries same drawbacks of other similar platforms. I'd rather use Hermes/Openclaw for oss or closed managed agents like Tasklet or Prajvis

qm has been a qemu (virtualization) command line tool probably for 20 years or so. The ignorance of these people inside the same field is egregious. (Oh, yes, there can be tools that are named the same in the same field, I get it)
Just want to thank the authors for the concise and consumable README.
Oh sweet, nice to see someone writing about multivalue databases.
Is it intended to support running on generic hosting?
> Learn your writing voice from past sends, then triage your inbox on a schedule — labels and reply drafts included

Oh grand. The world will be full of agents talking to agents and nothing getting done. So pretty much the same only I can just leave my machine awake to look busy.

i find something a bit funny in an ai project, written by ai, requiring human-written text with specific guidance to not use ai.

"Given that coding agents write most underlying code now, we'd prefer PRs in the form of human-written text. [...] Please do not have AI artificially expand what you'd like to do into a formal proposal."

i'm curious what this gains in practice and usability over standard chatbot slack integration.
Fake, boring, non-accomplishment.
not a very helpful title? maybe: "qm - a multiplayer agent harness for work"
The design concepts of “qm” and “Block Buzz” share some similarities in their underlying principles.
Other than the hosting providers, who has made money directly from running OpenClaw in constant loops?

This software appears to be yet another solution in search of a problem designed to burn as many tokens as possible.

"multiplayer"? Is this a game? I honestly don't know what that means in this context.
Sweet, now YC itself will be the startup :D.
this looks over engineered, gimmicky and lame. YC is clearly having a "crisis of meaning" in a post AI world
I have a simple question - Is QM another integrated development platform like VS Code?
This is smelling like an astroturfing campaign. This is a total nothing-burger
>We take contributions as human-written text, not code — see CONTRIBUTING.md. Describe the change you'd like informally in a .txt or .md file in adrs/, and if we're aligned we'll handle the implementation. Report vulnerabilities privately — see SECURITY.md, not a public issue.

Starting to think people were right when they talked about our industry itself having an AI psychosis problem.

Why does every online software say they are "multiplayer" now?

And then like 90% of the time there's nobody "there", and it's barely even collaborative.

If this is "multiplayer", does that make Chrome a Massively Multiplayer Online Browser?

Is everything a game?