back
user profile

padolsey

4,464karma·578submissions·June 29, 2011
about
padolsey.at.hn

I'm James. Living between Beijing and London. I like coding. Also plants. Stroke survivor & disability advocate. My dog's a whippet/iggy cross and is called Ducky. He's a beautiful lunatic.

* website: [j11y.io](https://j11y.io) * building: [nope](https://nope.net) * bsky: [@j11y.io](https://bsky.app/profile/j11y.io) * contact: https://tally.so/r/waYPvE * twitter: [@padolsey](https://x.com/padolsey) * book recommendations: [ablf.io](https://ablf.io) * me = founding eng @ [collective intelligence project](https://cip.org)

My work email is my first name at cip dot org.

---

![me](https://j11y.io/ducky.jpg)

---

My meet.hn token: meet.hn-10d164d4-bf7f-4431-898f-e02b34b836dc

recent activity (578 total)
comment
This is why I've found the safety research conducted by the likes of Anthropic and OAI to be so confusing. Like when they said that models are likely to blackmail developers in order to avoid bei…
11mo ago·view thread
comment
Hi jph! Not intending to subvert the thread, but I'd love to chat to someone like you. The non-profit I work at has been working on democratizing evals. This wouldn't be to ensure your in-h…
1y ago·view thread
comment
You're right; we should not even try. Better to have 0% compliance coverage and your honor than 90% coverage as best-effort.
1y ago·view thread
comment
> This is cool, but I’m a little skeptical. If Parachute uses AI agents to evaluate other models, who’s evaluating the AI agents? Usually you can run human-in-the-loop spot checks to ensure that th…
1y ago·view thread
comment
Most of these attacks succeed because app developers either don’t trust role boundaries or don’t understand them. They assume the model can’t reliably separate trusted instructions (system/develo…
1y ago·view thread
comment
Reminds me of haikus; to be true in nature, they must have a 'cutting word' to severely juxtapose, allowing two otherwise irreconcilable meanings to be bridged. A good haiku must be composed…
1y ago·view thread
comment
> But I will ask hard questions to see how well a candidate communicates in a stressful situation Well, stress fires off old cortex fight-or-flight response. It's like the worst possible test …
1y ago·view thread
comment
> Some companies genuinely care about this. Some even mention it in their job descriptions. They want candidates who perform well under pressure. Who are these self-important employers?? I mean, ou…
1y ago·view thread
comment
>> LLM's do not possess professional experience needed for successful therapy, such as knowing when to not say something as LLM's are not people. > Most people do not either. That a…
1y ago·view thread
comment
Entirely anecdotally ofc, I find that therapists often over-bias to formal diagnoses. This makes sense, but can mean the patient forms a kind of self-obsessive over-diagnostic meta mindset where every…
1y ago·view thread
comment
Ofc the title got its point across, but I'd argue to hold ourselves to higher standards of veracity. That's all.
1y ago·view thread
comment
Yep it's always a bit cute and funny, when you consider the absolute necessity of dopamine in basically every functionally relevant neural activation. Talk to a parkinson's patient about you…
1y ago·view thread
comment
> Using ChatGPT on cognitive tasks can reduce your brain connectivity by up to 50% and reduce your ability to recall information about the task by 8x. Argh people keep referencing this study as Gos…
1y ago·view thread
comment
Agree utterly. It's a real shame, and severely affects accessibility for disabled and elderly people. Not only UI discoverability but also the types of swiping or holding movements required on mo…
1y ago·view thread
comment
A know a person at the FCDO who had to routinely write letters in correspondance to those who'd sought Palmerston's advice on various matters. A hilarious internship.
1y ago·view thread
comment
I really recommend people study the measurement frailties and prompting sensitivities of LLM judges before employing them. They're valuable, but should be used with complete understanding of the …
1y ago·view thread
comment
Agreed! FWIW I am attempting to create an open-source wiki/watchdog eval platform -- weval.org -- , so we can all keep an eye on LLMs, their biases, and their general competencies without relyong…
1y ago·view thread
comment
How much transparency does Claude Code give you into what it's doing? I like IDE-integrated agents as they show diffs and allow focused prompting for specific areas of concern. And I get to contr…
1y ago·view thread
comment
IMO This is such disingenuous and misleading thing for Anthropic to have released as part of a system card. It'll be misunderstood. At best it's of cursory intrigue, but it should not be rea…
1y ago·view thread
comment
If we're paying for reasoning tokens, we should be able to have access to these, no? Seems reasonable enough to allow access, and then we can perhaps use our own streaming summarization models in…
1y ago·view thread
comment
> if you want to sculpt the kind of software that gets embedded in pacemakers and missile guidance systems and M1 tanks—you better throw that bot out the airlock and learn. But the bulk of us aren&…
1y ago·view thread
comment
I feel the same. On reflection, it's how I think I experience thoughts emerging in my head. Language gets derived from initially noisy embeddings. It's quite beautiful that we've ended …
1y ago·view thread
comment
Interesting! > People who don't pause exist more in their head than their body. The mind is top-down, rigid, quick, enforcing an established view. The mind is waiting for the other person to b…
1y ago·view thread
comment
> That's just an advertisement for their competitor. I think the simpler explanation is just laziness and no positive incentive or obligation, as opposed to proactive competitive practices..
1y ago·view thread
comment
I wonder if it's just our egos talking, telling us that for something to have value within this technical sphere it has to be complex and 'hard won'? I see this agent stuff as a pretty …
1y ago·view thread
comment
HTML was designed prior to XML fwiw, spinning off from SGML.
1y ago·view thread
comment
I wish they had some example completions in the paper and not just eval results. It would be really useful to see if there are any emergent linguistic tilts to the newly diverse responses...
1y ago·view thread