try a fresh coding session without skills, agents.md, system prompt and additional tools
I think you will be positively surprised how good current models like GPT 5.6 Sol are when they are not oversteered and context spammed
Here is a task (python templating) with 9 runs with OpenCode, Pi and smol
https://smolenv.com/t/nested-template-includes-60636/
you can read each run step by step and see what the agents are doing and how the system prompt and available tools are steering their behaviour to take longer and higher cost
(disclaimer: I'm working on smol)