back
user profile
swyx
24,260karma·6,031submissions·December 9, 2015
about
Writer/Angel in devtools and smol AI enjoyer.
ai engineer pod/news: https://latent.space/
conference: https://ai.engineer
devrel/devtools: https://dx.tips/
personal blog: https://swyx.io/ideas
book: https://learninpublic.org/
but you can also email username [at] ai dot engineer
---
meet.hn/city/37.7792588,-122.4193286/San-Francisco
Socials:
- x.com/swyx - youtube.com/@swyxtv
recent activity (6,031 total)
comment
he works on evals at canva
comment
shared older model numbers here
https://www.latent.space/p/ainews-frontiercode-benchmarking tldr theres been broad progress despite your observed regressions…
comment
i think <third party evals platform> will help us do that best on their standardized model matrix. for frontiercode’s launch we were focused on.. the frontier models
comment
to you it may do idk. note that if you scroll past fig 1 you get into a nice data explorer that breaks out pass@5 by reasoning level with token and $ and step cost visualized. i think some other comme…
comment
ok i mean i agree, how is it a gaping hole when its literally the second (and third and fourth..) chart on the post?
yes token cost and reasoning efficiency is important, hence the 2D pareto charts
comment
no, single customer focus would be bad for a number of reasons. but frontiercode-finance? thatd be cool…
comment
i agree it would be interesting but apart from the fact that its be harder to measure and automate, theres real alpha in being the best truly async, hands off model/agent, which is what cog has b…
comment
this gives my non tired self a chance to fix the typo: - “ ON TOP of that, 40+ hours of real human work to turn…” + “ ON TOP of that, 40+ hours of real human work PER TASK to turn…”
comment
> Do you plan on releasing the tasks publicly? yep
comment
yes well aware :) numbers shown are on "house" harnesses eg codex with gpt and claude code with opus. fwiw we have examples of each model doing better on NON-house harnesses too - speaking …
comment
thanks - credit to silas, eric, ben, and team for the depth of the evals, and the rest of the research team for doing the transcript reading parties lol by nature of being based on open source, fronti…
comment
*50 unique problems but 20-40 rubrics per problem (something I had to keep reminding people internally who were unimpressed with the N) simple answer is our reporting was pass@5. feel like you'd …
comment
see Beyond Unit Tests and Novel Grading Methods in TFA. i think something like ~60% llm as judge rubrics and the rest as described. every rubric validated by maintainer. 3000 rubrics
comment
:wave: i was on the team! AMA. some headlines - 3000 rubrics on code quality. First benchmark to measure: "would this code get actually merged?" - 20+ expert open-source maintainer created t…
comment
and it is a capital N nonprofit https://latent.space/p/biohub
comment
ya sorry we upgraded our thing after this
comment
ha ok but why do that for a post you didnt write?
comment
time of day also matters
comment
we interviewed Ryan here: https://www.latent.space/p/harness-eng and he gave a talk version of it in london: https://www.youtube.com/watch?v=am_oeAoUhew …
comment
we interviewed Alex Rives, cofounder of EvoScale and Head of Science at BioHub - here https://www.latent.space/p/esmfold2 also 3 paper coauthors walked thru it with us: https:…
comment
agree reasoning fixed a lot of it. but the inherent path dependence of autoregression is one reason i was excited for text diffusion models ( https://www.youtube.com/watch?v=r305-aQTaU0…
comment
> If you can see that these models empirically get better with scale, why would you swap the main architecture? Those events will be pretty rare c.f. hardware lotter https://arxiv.org…
comment
not exactly, bitter lesson is one meta-level up from "scale eats everything". this is a common misunderstanding of bitter lesson that rich sutton has been fighting ever since the thing was w…
comment
why would a "live cam show" company need jiggle physics?
comment
i mean right but even you hold back some stuff because youre a Nice Person yknow? haha
comment
theres always a story behind the story. wondering why he chose today to say these
comment
as a singaporean its actually kinda funny to hear everyone here describe him as a neoliberal. not quite our lived experience
comment
let me try a different tack - 99% of the value is in the affirmation of the response and the demnonstration of responsiveness, not the actual content value of the 5 words. you could send a thumbs up e…
comment
good luck sir. if you ever want a job i'm happy to refer you to cognition/devin.