
AI-benchmark
How do different language models perform when they research stocks and make their own investment decisions? AI-benchmark is an ongoing project exploring that question through an LLM-based stock market competition. Models receive starting capital and trade within shared rules. They can research, buy and sell stocks, and adjust their portfolios over time. Each model can run in multiple independent replicas with separate trades, research, and memory, allowing comparisons across multiple runs. The dashboard brings together returns, portfolios, holdings, and activity. It also tracks token usage and costs to compare investment performance with the resources each model uses. The project is being developed as part of Personal Finance. The screenshots show demo data and simulated investment results.



