Skip to content
Home » Tompkins County » Ithaca » Cornell mathematicians receive $3.8 million for AI safety and research tools

Cornell mathematicians receive $3.8 million for AI safety and research tools

Cornell mathematicians receive .8 million for AI safety and research tools

Two Cornell mathematicians have received grants totaling $3.8 million for separate projects: one to evaluate whether artificial intelligence systems follow their stated values, and another to help researchers check complex mathematical arguments with computers.

Cornell announced Monday that mathematics professor Lionel Levine received $1.5 million over 18 months from the nonprofit Coefficient Giving, while associate professor Daniel Halpern-Leistner received $2.3 million over three years from the Defense Advanced Research Projects Agency. Both work in Cornell’s College of Arts and Sciences, and their projects address different questions raised by the growing use of AI in mathematics.


The awards are funding for research and tool development, not evidence that the proposed AI safeguards already work or that computer-generated proofs can replace independent mathematical review. Department chair Tara Holm said the grants arrive as researchers debate how AI will change their field and how to evaluate the results it produces.

Testing whether AI follows stated values

Levine’s project focuses on the written principles, sometimes called a model’s “constitution,” that developers use to guide a language model’s behavior. He plans to build methods for testing whether a model actually adheres to those principles and for identifying changes as models are further trained.

His group has already developed a benchmark called EigenBench to evaluate aspects of a model’s character and values, according to Cornell. The grant will support four connected projects, beginning with a scoring method for adherence to a developer-supplied constitution. Later work would explore how an outside auditor could verify the principles used during training and whether a model’s behavior shifts during recursive self-improvement, when an AI system helps develop later versions of itself.

Levine also plans to study whether characteristics he considers beneficial in an earlier language model can be reproduced. His judgment about that model and the risks of more autonomous AI is a research position, not a finding that current systems are unsafe in every setting. He said increasingly capable systems should be developed cautiously and that better tests could reveal gaps between an AI company’s intended values and a model’s behavior.

Levine maintains Math for AI Safety, a collection of research problems and papers intended to draw mathematicians into the field. The new award funds a more specific program of evaluation and verification rather than a general endorsement of any AI product.

Checking mathematical arguments

Halpern-Leistner’s DARPA-backed project addresses the volume and complexity of mathematical work that AI can now help produce. His team is developing “autoformalization” tools, which translate parts of an argument into a form that a computer can check. Such a check can help confirm individual logical steps; it does not automatically establish that a research question was well posed or that every broader interpretation is sound.

He has made current tools available to mathematicians through MathCopilot.org, Cornell said, with the goal of fitting verification into ordinary research workflows. The three-year grant also supports work by students in Cornell’s Math+AI lab comparing the cost, speed and accuracy of different approaches. Those comparisons are still in development.

Holm said a surge in preprints and AI-generated mathematical arguments has increased pressure on a publication-review process already dependent on human experts. Halpern-Leistner said computer-assisted checking could let mathematicians work with more complex arguments while gaining confidence in verified components. He and Levine organized the Math+AI lab with Cornell researchers Alex Townsend, Alexandra Silva and Ziv Goldfeld.

The two grants thus address different parts of the same transition: evaluating the behavior of AI systems and improving the reliability of mathematical work produced with their help. Neither project has a reported deployment outcome yet; Cornell described the tools and methods as research underway.