Skip to main content

Laying the foundation for AI-powered discovery

Northwestern researcher Chris Wolverton helped develop today’s AI-driven approach to materials discovery.

Chris Wolverton’s research group at Northwestern University uses computational tools and artificial intelligence to tackle today’s most pressing challenges in energy and materials science. For more than a decade, the group has played a critical role in shaping the use of AI for materials discovery, helping to reduce the time, cost, and material waste involved in developing cleaner, more efficient, and longer-lasting technologies.

In late 2010, Wolverton’s group began testing the university’s new Quest supercomputer by running large numbers of materials calculations. As the researchers improved the process and automated the calculations, the project grew beyond its original purpose. By 2013, the group had built a database containing predicted properties for thousands of materials. Wolverton, currently the Frank C. Engelhart Professor of Materials Science and Engineering, and his team had laid the foundation for a powerful resource for AI-assisted materials discovery.

Three years later, that work resulted in a paper published in JOM introducing the Open Quantum Materials Database (OQMD). Created by James Saal, Scott Kirklin, Muratahan Aykol, Bryce Meredig, and Chris Wolverton, OQMD was designed as an open resource that lets researchers quickly search and compare the predicted properties of hundreds of thousands of materials. A follow-up study published in npj Computational Materials in 2015 showed that the database's predictions closely matched experimental results, helping to establish OQMD as a trusted resource for materials scientists worldwide.

What set OQMD apart was that it looked beyond materials scientists had already discovered. In addition to known materials, the database included hundreds of thousands of computer-generated candidates based on existing crystal structures. That gave researchers a way to search for promising new materials that had never been studied before, opening up far more possibilities for discovery.

OQMD changed how many researchers approached materials discovery. Before it existed, research groups usually ran their own calculations, kept the results on local servers, and rarely shared the data. That meant scientists often spent time repeating work that others had already done. Without a large, publicly available database for comparison, identifying promising new materials was also a slow process that could take years before leading to real-world applications.

When the Obama administration launched the Materials Genome Initiative in 2011, investing more than $400 million to speed up the discovery and redevelopment of new materials, the Materials Project and OQMD were among the initiative’s first major open databases. These resources gave researchers everywhere free access to a growing collection of materials data. The openness was intentional. As Wolverton later explained, “Our philosophy from the very beginning was that our database should be open.”

The database continued to expand. By the end of 2014, OQMD included nearly 286,000 materials. It passed one million entries in 2021 and today contains more than 1.4 million materials that researchers around the world can search and download for free. To celebrate the millionth entry, the team honored the database’s original authors by translating their last names into chemical elements and combining them into a real compound, KMgAlWS₆.

The 2013 paper has since been cited more than 3190 times, making it JOM's most-cited paper. The 2015 follow-up paper has been cited more than 2,500 times and is the most cited non-review paper in npj Computational Materials. Today, researchers around the world use OQMD to study everything from batteries and solar fuels to structural materials and thermoelectrics, making it one of the most widely used resources in computational materials science.

Wolverton’s group recognized early that OQMD could be more than a searchable catalog. Because it contained a large, consistently calculated set of materials properties, it could also provide the training data for machine learning. In a pioneering 2014 study, Bryce Meredig, a former PhD student in Wolverton’s research group, Wolverton, and colleagues used data from OQMD to train a model that could rapidly estimate the formation energies of new compositions using about one-millionth of the computing time required for the underlying quantum-mechanical calculations. The team used the model to screen roughly 1.6 million candidate compositions and predict 4,500 potentially stable materials. The work was among the early demonstrations that machine learning could learn patterns from a large materials database and use them to explore chemical possibilities far beyond the entries already calculated.

The team expanded that idea in 2016, when Logan Ward, then a PhD student in Wolverton’s research group, led the development of a general-purpose machine-learning framework that used OQMD data to predict materials properties and identify promising new materials. The researchers introduced a general-purpose machine-learning framework for predicting materials properties. Trained on 228,676 compounds from OQMD, the models predicted formation energy, band gap, and volume, then used those predictions to identify promising solar-cell materials.

This shift from using computation to populate a database to using the database to train predictive models helped usher in today’s AI-driven approach to materials discovery. Instead of performing an expensive quantum-mechanical calculation for every candidate, researchers could use machine learning to screen vast numbers of possibilities quickly and reserve detailed calculations and experiments for the most promising ones.

OQMD’s impact has reached far beyond academia. Google DeepMind used data from OQMD and the Materials Project to help train early versions of GNoME, its AI system for discovering new materials. Two of OQMD’s original contributors, Muratahan Aykol and Vinay Hegde, went on to work at Google DeepMind, helping carry the ideas behind OQMD into the next generation of AI-powered materials discovery. Aykol is now at Periodic Labs, a company that aims to create an AI scientist.

Alumni of the Wolverton research group are also applying the team’s research to real-world industrial applications through entrepreneurship. In 2013, Bryce Meredig, who earned his PhD in Wolverton's research group the previous year, co-founded Citrine Informatics with Greg Mulholland to bring data-driven materials discovery into industry. James Saal, a former postdoctoral researcher in Wolverton’s lab, later joined the company. Today, Citrine works with manufacturers including Panasonic, Michelin, Rolls Royce, and LANXESS and has raised more than $80 million in funding.

OQMD’s influence has continued to grow. Meta’s Open Materials 2024 dataset cites it as one of the projects that helped shape its development, and researchers regularly point to OQMD alongside the Materials Project and AFLOW as one of the field's foundational open databases. Its impact also continues in current research. A 2026 study of sodium-ion batteries, for example, combined data from OQMD, the Materials Project, AFLOW, and GNoME, integrating several generations of materials databases and AI tools into a single discovery pipeline.

Many of the ideas behind OQMD are now standard practice. Today, researchers use AI, large open databases, and automated laboratories to search for promising new materials much faster than was once possible. Results from those experiments are then fed back into shared databases, creating a cycle of discovery built on the same open-data principles that helped make OQMD so influential.

Wolverton continues to build on that work today, using computational modeling and machine learning to develop new materials for batteries, clean energy technologies, and other energy applications. In 2025, he received the Materials Research Society's Materials Theory Award in recognition of his contributions to the field.

What began as a way to test a new supercomputer grew into one of the world's most widely used materials databases. Long before today's AI systems could search for new materials, researchers first had to build the datasets those systems learn from. OQMD helped lay that foundation.

Selected Publications

Shen, J., Griesemer, S.D., Gopakumar, A., Baldassarri, B., Saal, J.E., Aykol, M., Hegde, V.I., & Wolverton, C. (2022). Reflections on one million compounds in the open quantum materials database (OQMD). Journal of Physics: Materials, 5(3), 031001.

Meredig, B., Agrawal, A., Kirklin, S., Saal, J.E., Doak, J.W., Thompson, A., Zhang, K., Choudhary, A., & Wolverton, C. (2014). Combinatorial screening for new materials in unconstrained composition space with machine learning. Physical Review B, 89(9), 094104.

Ward, L., Agrawal, A., Choudhary, A., & Wolverton, C. (2016). A general-purpose machine learning framework for predicting properties of inorganic materials. npj Computational Materials, 2, 16028.

Saal, J.E., Kirklin, S., Aykol, M., Meredig, B., & Wolverton, C. (2013). Materials Design and Discovery with High-Throughput Density Functional Theory: The Open Quantum Materials Database (OQMD). JOM, 65(11), 1501–1509.

Kirklin, S., Saal, J.E., Meredig, B., Thompson, A., Doak, J.W., Aykol, M., Rühl, S., & Wolverton, C. (2015). The Open Quantum Materials Database (OQMD): assessing the accuracy of DFT formation energies. npj Computational Materials, 1, 15010.