An AI model can explain an earnings report, calculate financial ratios and produce a convincing investment thesis. For investors, the next question is whether any of this improves the decisions we make.
Does it help identify companies whose outlook is improving? Can it distinguish a genuine upgrade from reassuring management language? And does the information it extracts tell us anything about returns that the market has not already priced in?
I use financial reasoning benchmarks as an initial screen for accuracy and cost. They help shortlist a model; the investment test comes next.
My starting point is a narrow research question: can changes in a company’s outlook improve a simple stock-selection process?
I want a rule that is transparent enough to explain, repeatable enough to test, and practical enough to trade after the information becomes available. The language model has a specific role within that process: extracting comparable forward guidance from earnings releases.
Earnings announcements offer a useful setting. They bring together reported numbers, management expectations and changes in the business. Comparing these across reporting periods gives us a way to turn narrative information into a measurable signal.
Here I develop the scoring rule, work through an example and specify the test I would use before risking capital.


