Evaluating Large Language Models in Scientific Discovery (arXiv:2512.15567) — a large multi-institution effort led by Zhangde Song, asking whether LLMs can actually run the loop of hypothesis, observation, and revision, rather than recite decontextualized facts.