back
user profile
dhorthy
1,124karma·268submissions·March 2, 2018
about
building @ humanlayer.com
recent activity (268 total)
comment
they seem to be able to do the call-stack diff stuff pretty well still
comment
so wait is the finding that most of those skills reduce pass rates against SCB? wild
comment
why would that worry you?
comment
agree, i think the implication is that low quality code is harder to change in the future
comment
somebody get this man a curl-pipe-bash stat
comment
oh i really like the idea of flipping around the order of checkpoints and comparing results. Could be an interesting way to increase/decrease difficulty even i will look into how easy it would be…
comment
Yeah I would hold that models don’t know how to simplify because most rl/benchmarks doesn’t penalize complexity
comment
Yeah the main reason I skipped fable was because we have a ZDR with anthropic and I didn’t feel like spinning up another account to circumvent that. Next run will have fable and sol
comment
my issue with frontier code is that it uses a model judge for quality whereas slop code bench forces a model to grapple with its own garbage code in order to receive a functionality reward
comment
i laughed at the pelican bit its good yes the labs will always prioritize the vibeslop dopamine casino as far as I can tell - making the models useful and addictive for unsophisticated users, sometime…
comment
yeah someone will have to re-run this bench on various effort levels. unfortunately it is not cheap
comment
I agree this is an option, and the next thing on my radar is to try with a more realistic "factory-shaped" harness where you have feedback from linters and other models after each coding epi…
comment
> - what "maintainable" is is probably some high dimensional space described by these signals; it'd probably require some human labeling to figure out where this space is this is a n…
comment
yes sol is still my daily driver for most coding tasks I did find opus 5 quite handy for general knowledge work and visual design, without the cost of fable (e.g. the graphics in this post are made by…
comment
no i'm spinning those up at some point this week. here's the first few prompts I used (claude opus 5 as the research orchestrator), (these were interspersed with lots of tools and assistant …
comment
i hope that is because you hate slop and not because you write it
comment
yeah this was just a start - the fastest cheapest thing we could try for a brand new model. I'm hoping to do some more work with sol/fable in the mix as well as exploring more languages and …
comment
the one good thing about the current ios version is that it is the best it will ever be from this point forward. all future versions will be worse
comment
I agree this rocks I do this almost daily
comment
yes well said
comment
don't forget the andon cord
comment
i humbly disagree - horizontal means touching one plane of the stack across, vertical means cutting down through it and touching multiple layers https://en.wikipedia.org/wiki/Vert…
comment
i did a write up on fable while it was out - it can do big refactors, but it does not know what to change without human steering. For that, you need humans to know what to ask for. lights off is still…
comment
its the optimizing utilization instead of overall throughput all over again. eli goldratt talked about this in the 1970s[1]. we still haven't learned 1 - https://en.wikipedia.org/…
comment
exactly. a great pr is a joy to review. we've found some success in agents generating static HTML walkthroughs that order the diffs in something other than GitHub's default alphanumeric orde…
comment
i guess to clarify my contention: generic prompting like "review this code" or "make the architecture better" will raise the floor but cannot come close to human-quality code witho…