To date, it has solved 65 PutnamBench problems during development.
LLMs tested via API: GPT-5.2,GPT-5.6 Luna-Pro, DeepSeek-V4-Flash, DeepSeek-V4-Pro, and Qwen3.7-Max.
Development is ongoing. Would love to see others use it to attempt unsolved problems in parallel.