back
user profile

swyx

24,260karma·6,031submissions·December 9, 2015
about
Writer/Angel in devtools and smol AI enjoyer.

ai engineer pod/news: https://latent.space/

conference: https://ai.engineer

devrel/devtools: https://dx.tips/

personal blog: https://swyx.io/ideas

book: https://learninpublic.org/

but you can also email username [at] ai dot engineer

---

meet.hn/city/37.7792588,-122.4193286/San-Francisco

Socials:

- x.com/swyx - youtube.com/@swyxtv

recent activity (6,031 total)
comment
he works on evals at canva
2mo ago·view thread
comment
shared older model numbers here https://www.latent.space/p/ainews-frontiercode-benchmarking tldr theres been broad progress despite your observed regressions…
2mo ago·view thread
comment
i think <third party evals platform> will help us do that best on their standardized model matrix. for frontiercode’s launch we were focused on.. the frontier models
2mo ago·view thread
comment
to you it may do idk. note that if you scroll past fig 1 you get into a nice data explorer that breaks out pass@5 by reasoning level with token and $ and step cost visualized. i think some other comme…
2mo ago·view thread
comment
ok i mean i agree, how is it a gaping hole when its literally the second (and third and fourth..) chart on the post? yes token cost and reasoning efficiency is important, hence the 2D pareto charts
2mo ago·view thread
comment
no, single customer focus would be bad for a number of reasons. but frontiercode-finance? thatd be cool…
2mo ago·view thread
comment
i agree it would be interesting but apart from the fact that its be harder to measure and automate, theres real alpha in being the best truly async, hands off model/agent, which is what cog has b…
2mo ago·view thread
comment
this gives my non tired self a chance to fix the typo: - “ ON TOP of that, 40+ hours of real human work to turn…” + “ ON TOP of that, 40+ hours of real human work PER TASK to turn…”
2mo ago·view thread
comment
> Do you plan on releasing the tasks publicly? yep
2mo ago·view thread
comment
yes well aware :) numbers shown are on "house" harnesses eg codex with gpt and claude code with opus. fwiw we have examples of each model doing better on NON-house harnesses too - speaking …
2mo ago·view thread
comment
thanks - credit to silas, eric, ben, and team for the depth of the evals, and the rest of the research team for doing the transcript reading parties lol by nature of being based on open source, fronti…
2mo ago·view thread
comment
*50 unique problems but 20-40 rubrics per problem (something I had to keep reminding people internally who were unimpressed with the N) simple answer is our reporting was pass@5. feel like you'd …
2mo ago·view thread
comment
see Beyond Unit Tests and Novel Grading Methods in TFA. i think something like ~60% llm as judge rubrics and the rest as described. every rubric validated by maintainer. 3000 rubrics
2mo ago·view thread
comment
:wave: i was on the team! AMA. some headlines - 3000 rubrics on code quality. First benchmark to measure: "would this code get actually merged?" - 20+ expert open-source maintainer created t…
2mo ago·view thread
comment
and it is a capital N nonprofit https://latent.space/p/biohub
2mo ago·view thread
comment
ya sorry we upgraded our thing after this
2mo ago·view thread
comment
ha ok but why do that for a post you didnt write?
2mo ago·view thread
comment
time of day also matters
2mo ago·view thread
comment
we interviewed Ryan here: https://www.latent.space/p/harness-eng and he gave a talk version of it in london: https://www.youtube.com/watch?v=am_oeAoUhew …
2mo ago·view thread
comment
we interviewed Alex Rives, cofounder of EvoScale and Head of Science at BioHub - here https://www.latent.space/p/esmfold2 also 3 paper coauthors walked thru it with us: https:&#…
2mo ago·view thread
comment
agree reasoning fixed a lot of it. but the inherent path dependence of autoregression is one reason i was excited for text diffusion models ( https://www.youtube.com/watch?v=r305-aQTaU0…
2mo ago·view thread
comment
> If you can see that these models empirically get better with scale, why would you swap the main architecture? Those events will be pretty rare c.f. hardware lotter https://arxiv.org&#x…
2mo ago·view thread
comment
not exactly, bitter lesson is one meta-level up from "scale eats everything". this is a common misunderstanding of bitter lesson that rich sutton has been fighting ever since the thing was w…
2mo ago·view thread
comment
why would a "live cam show" company need jiggle physics?
2mo ago·view thread
comment
i mean right but even you hold back some stuff because youre a Nice Person yknow? haha
2mo ago·view thread
comment
theres always a story behind the story. wondering why he chose today to say these
2mo ago·view thread
comment
as a singaporean its actually kinda funny to hear everyone here describe him as a neoliberal. not quite our lived experience
2mo ago·view thread
comment
let me try a different tack - 99% of the value is in the affirmation of the response and the demnonstration of responsiveness, not the actual content value of the 5 words. you could send a thumbs up e…
2mo ago·view thread
comment
good luck sir. if you ever want a job i'm happy to refer you to cognition/devin.
2mo ago·view thread