This makes me wonder too about the entire premise and worthiness of these evals. They orient themselves around normal one-shot interactions with a likely non-sys-prompted model with no built up context or memory of the person. I doubt the mentioned 'job loss' scenario is even contextually seen as a 'loss'; it is only a circumstance descriptor, a single snapshot without a history. Maybe to get the best advice we actually need to tell the LLM our entire story, not just a narrow request for a question; a question that - itself - is biased to our own imaginings of what problem we perceive ourselves as having, which humans are often bad at.
First, from a technical standpoint the required context window would be massive if you're looking at a person's career/life holistically. Probably solvable, but definitely something to be aware of.
Second, privacy goes completely out the window since you're sharing everything. You don't know what's relevant and what's not up front so you need to provide everything.
Third, you would need a training dataset of all those input variables and their outcomes to be able to provide any sort of useful output. The first set of people to share everything wouldn't be able to derive any value from the tool, and I think you'd be hard pressed to convince enough people to do it to get a useful dataset.
You don't necessarily need to provide everything. Arthur (our AI) is smart enough to see exactly which information it needs to answer a given question. but, yes, the more information you provide the easier of a time the AI will have in answering your question. Arthur doesn't guess. if there is crucial information it needs he will ask for it. it doesn't have to be a Plaid hook up, a csv or even a simple user response is a start.
On your third point — you'd need an outcomes dataset — that's true for traditional ML, but it's not how this works. The normative layer is finance itself (life-cycle theory, tax rules, amortization) implemented as deterministic calculators, with the LLM doing explanation and elicitation. The paper under discussion is sort of the proof: the models already give theory-aligned advice with zero outcome training. The gap it found is input quality and statelessness, not a missing training set.
Why would it be massive? The application layer typically compacts a profile of information about the users financial situation when offered. I doubt many of us have financial situations that would exceed the context window.
> Third, you would need a training dataset of all those input variables and their outcomes to be able to provide any sort of useful output. The first set of people to share everything wouldn't be able to derive any value from the tool, and I think you'd be hard pressed to convince enough people to do it to get a useful dataset.
Would you 'need' a training dataset of input variables and their outcomes for an LLM? Certainly for traditional ML, but the LLM toolcalling can simulate what an astute user should statistically do in their situation based on information on the internet and reason about the different constraints.
Like, do you have $1,000 in an emergency fund? No? Start there.
The more information you give pendragon the better your answers will be. Pendragon also produces a history of decisions, memories, and plans so it can keep you on track and have a better understanding of your overall financial health. Pendragon helps you achieve your goals by providing detailed financial advise and specific actions you can take to achieve whichever goal you have.
I’d rather get no response than be patronized, so no linking to corporate policy documents on your website please.
I know you said that you don't want a TOS link, here's something better. This is our constitution. https://pendragon.foxtrotcommunications.net/constitution
I worry about my data being sold if I used a tool like this, but I should probably be more worried about my actual credit card purchase data being sold (because it is).