Engineering PapersSearch

SEARCH · Engineering Papers

Results for “open data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Indicators of Global Climate Change 2023: annual update of key indicators of the state of the climate system and human influence

Intergovernmental Panel on Climate Change (IPCC) assessments are the trusted source of scientific evidence for climate negotiations taking place under the United Nations Framework Convention on Climate Change (UNFCCC). Evidence-based decision-making needs to be informed by up-to-date and timely information on key indicators of the state of the climate system and of the human influence on the global climate system. However, successive IPCC reports are published at intervals of 5–10 years, creating potential for an information gap between report cycles. We follow methods as close as possible to those used in the IPCC Sixth Assessment Report (AR6) Working Group One (WGI) report. We compile monitoring datasets to produce estimates for key climate indicators related to forcing of the climate system: emissions of greenhouse gases and short-lived climate forcers, greenhouse gas concentrations, radiative forcing, the Earth's energy imbalance, surface temperature changes, warming attributed to human activities, the remaining carbon budget, and estimates of global temperature extremes. The purpose of this effort, grounded in an open-data, open-science approach, is to make annually updated reliable global climate indicators available in the public domain. As they are traceable to IPCC report methods, they can be trusted by all parties involved in UNFCCC negotiations and help convey wider understanding of the latest knowledge of the climate system and its direction of travel. The indicators show that, for the 2014–2023 decade average, observed warming was 1.19 [1.06 to 1.30] °C, of which 1.19 [1.0 to 1.4] °C was human-induced. For the single-year average, human-induced warming reached 1.31 [1.1 to 1.7] °C in 2023 relative to 1850–1900. The best estimate is below the 2023-observed warming record of 1.43 [1.32 to 1.53] °C, indicating a substantial contribution of internal variability in the 2023 record. Human-induced warming has been increasing at a rate that is unprecedented in the instrumental record, reaching 0.26 [0.2–0.4] °C per decade over 2014–2023. This high rate of warming is caused by a combination of net greenhouse gas emissions being at a persistent high of 53±5.4 Gt CO 2 e yr -1 over the last decade, as well as reductions in the strength of aerosol cooling. Despite this, there is evidence that the rate of increase in CO 2 emissions over the last decade has slowed compared to the 2000s, and depending on societal choices, a continued series of these annual updates over the critical 2020s decade could track a change of direction for some of the indicators presented here.

54 ENVIRONMENTAL SCIENCES

An open-source data storage and visualization platform for collaborative qubit control

Developing collaborative research platforms for quantum bit control is crucial for driving innovation in the field, as they enable the exchange of ideas, data, and implementation to achieve more impactful outcomes. Furthermore, considering the high costs associated with quantum experimental setups, collaborative environments are vital for maximizing resource utilization efficiently. However, the lack of dedicated data management platforms presents a significant obstacle to progress, highlighting the necessity for essential assistive tools tailored for this purpose. Current qubit control systems are unable to handle complicated management of extensive calibration data and do not support effectively visualizing intricate quantum experiment outcomes. In this paper, we introduce Qubit Control Storage and Visualization ( QubiCSV ), a platform specifically designed to meet the demands of quantum computing research, focusing on the storage and analysis of calibration and characterization data in qubit control systems. As an open-source tool, QubiCSV facilitates efficient data management of quantum computing, providing data versioning capabilities for data storage and allowing researchers and programmers to interact with qubits in real time. The insightful visualization are developed to interpret complex quantum experiments and optimize qubit performance. QubiCSV not only streamlines the handling of qubit control system data but also improves the user experience with intuitive visualization features, making it a valuable asset for researchers in the quantum computing domain.

97 MATHEMATICS AND COMPUTING

Using AI to Reproduce Neutrino Cross Section Analysis - Prototyping the Neutrino Discovery Platform

The Neutrino Discovery Platform (NDP) aims to accelerate DUNE-era science by making the neutrino program's existing datasets analyzable through fast, reproducible, and auditable workflows. We report a working version of two of its layers, data curation and agentic orchestration, built and tested end to end on MINERvA open data. The guiding lesson throughout is that a cross section is a measurement, and not just a plotted shape, only if it carries a defensible systematic-uncertainty budget, a trustworthy unfolding, and a reproducible record. Using a single medium-energy playlist pair from the MINERvA open-data release (about $2.05\times10^{17}$ protons on target of data), we first reproduced the shapes of two published charged-current inclusive $\nu_\mu$ measurements through a complete extraction ladder: selection, background subtraction, D'Agostini unfolding, efficiency correction, and flux normalization. These shape-level reproductions ran and tracked the published results, but they lacked the systematic-uncertainty machinery that defines a MINERvA cross section. To supply it, we vendored and built the MINERvA Analysis Toolkit and developed a many-universe systematic-uncertainty tool that produces a portable covariance artifact, a parallel event-loop runner, and a per-run auditability harness. Validated against a published covariance release, the toolchain reproduces the released statistical, flux, and muon-energy-scale terms and shows that they account for roughly 63\% of the total variance, with the remainder unreleased. Using this same infrastructure, we then performed a measurement of our own design, the hadronic recoil-energy distribution of low-energy ($E_\nu<2.5$~GeV) charged-current inclusive events, and found data/simulation shape agreement of $\chi^2/\mathrm{ndf}=1.26$. Together these results show that the platform supports original physics and not only reproductions.

Breaux, Auto [Tulane U. (main)]

Open-Source Data Analysis Tool for Spectral Small-Angle X-ray Scattering Using Spectroscopic Photon-Counting Detector

Spectral small-angle X-ray scattering (sSAXS) is a powerful technique for material characterization from thicker samples by capturing elastic X-ray scattering data in angle- and energy-dispersive modes at small angles. This approach is enabled by the use of a 2D spectroscopic photon-counting detector that provides energy and position information of scattered photons when a sample is irradiated by a polychromatic X-ray beam. Here, we describe an open-source tool with a graphical interface for analyzing sSAXS data obtained from a 2D spectroscopic photon-counting detector with a large number of energy bins. The tool takes system geometry parameters and raw detector data to output 1D scattering patterns and a 2D spatially-resolved scattering map in the energy range of interest. We validated these features using data from samples of caffeine powder with well-known scattering peaks. This open-source tool will facilitate sSAXS data analysis for various material characterization applications.

Chemistry

Multi-system analysis of offshore geologic carbon storage: a review of open-source data science solutions

Geologic carbon storage projects are maturing worldwide and the footprint of deployment in the offshore is expanding. At present, there are ten projects in operation or that have been completed, more than 50 in construction and development, and dozens of characterization studies completed or underway. Offshore geologic carbon storage offers potential benefits over onshore geologic carbon storage. These offshore projects are generally remote in location, distant from population centers, and avoid complicated pore space rights while having abundant prospective storage potential. Some offshore fields targeted for carbon storage have comparatively fewer prior borehole penetrations except for areas that have been explored for petroleum production, minimizing potential issues such as pressure interference and infrastructure impacts. Yet offshore geologic carbon storage projects face distinctive technical and economic challenges, such as seafloor geohazards (e.g., seabed instability), expensive maritime transport, and meteorological-oceanographic conditions that can damage infrastructure and impact operations. Analytical capabilities and improved computational speeds have advanced engineering, earth and energy sciences in the wake of the arrival of modern data science over the last decade. These advancements have created an opportunity for integrated, multi-systems modeling approaches utilizing artificial intelligence and machine learning that are no longer limited by computational issues. Analytical tools developed alongside this advancement in data science can be leveraged to calibrate the potential advantages and challenges of carbon storage operations in the offshore. New methods and approaches that incorporate data science to analyze multiple aspects of engineered and natural systems can provide insights that complement the characterization and onsite engineering that traditional commercial and operational software addresses. These new methods and approaches can potentially improve the outcome of energy operations and carbon storage. Providing multi-system, science-driven data analytics enhances the knowledge base that offshore developers, operators, and regulatory bodies may draw from to improve offshore site selection and operational efficiency. Here, we provide a brief synopsis of geologic carbon storage efforts to date, an overview of the engineered and natural systems involved in offshore geologic carbon storage, and a review of publicly available, open-source, offshore and/or carbon storage related data- and science-driven tools developed by 2010 or later that are suitable for screening and assessing regions for offshore geologic carbon storage.

artificial intelligence

Initial Mobility Analysis for ORNL VA-EDH Synthetic Populations

Travel burdens are a major barrier to healthcare access among US Veteran patient populations, particularly those residing in rural areas. Spatial accessibility to points of care for US Veteran populations is commonly assessed in two ways. The first approach uses open data from the US Census to represent collective travel burdens, for example the distance between population-weighted census tract centroids and VHA points of care. The second approach uses restricted-access VHA patient data to measure travel costs (e.g., distance, time) for accessing points of care with respect to geolocated patient addresses and real or approximated transportation networks. While the advantage of the open data approach lies in its reproducibility, it has notable limitations in its tendency to infer individual travel behavior from aggregate population characteristics, a problem known as ecological fallacy. Conversely, while the patient data approach is able to account for individual travel behavior, its ability to account for localized access disparities (e.g., a neighborhood with exceptionally high transportation costs) and patient demographics is limited as protecting individual patient data requires their storage in closed systems with limited capacity for adequately modeling real-world travel patterns or for supplementing patient attributes. Additionally, the patient data approach cannot account for veterans who are not enrolled in the VHA system but who may be eligible for care. These challenges limit the ability to perform “what if” analyses on the effects of place-specific interventions on veteran populations with high access barriers to healthcare. To address these challenges, we explore the application of realistic synthetic populations to examine travel burdens and spatial accessibility issues among veteran patient populations. Synthetic populations provide a virtual, individually-resolved and cross-sectional representation of the veteran patient population that enables investigation of spatial access to points of care in ways in which aggregate data and patient data do not. First, synthetic populations allow one to directly assess how individuals access points of care, from synthesized residential locations to outpatient facilities on real-world transportation networks. Modeling access to points of care at the individual scale addresses the ecological fallacy problem associated with using aggregated census data to represent veteran populations and patterns of movement. Second, synthetic populations provide a means of completely representing an area’s veteran population using only publicly available, anonymized census microdata from the American Community Survey (ACS) to ensure the privacy of real-world individuals. Generating synthetic populations from the ACS also expands descriptive characteristics beyond what patient data typically offers to include socio-demographic, economic, housing, and mobility attributes. More detailed profiles of both VHA patient populations and veterans not enrolled in the VA system will provide a comprehensive picture of groups that may benefit from interventions or outreach. As an initial exercise for using synthetic populations to measure veteran travel burdens to VA care, we apply Oak Ridge National Laboratory’s (ORNL) UrbanPop capability to generate a series of synthetic VHA patient populations for 9 Veterans Integrated Services Networks (VISN) market areas in 9 Census Divisions across the continental United States, which are listed in Table 1. We use UrbanPop to produce synthetic populations for the VISN markets selected for each US Census Division, then assign VA outpatient clinic destinations to synthetic VHA patients based on travel about each VISN market’s road network. To demonstrate using the synthetic populations to evaluate healthcare travel burdens, we compare the time-based impedance between simulated home locations and VA outpatient clinics in each VISN market. We then perform validation exercises on the synthetic populations with respect to neighborhood (block group) demographic composition as well as patient mobility, comparing aggregate origin-destination statistics for the synthetic population to outpatient visits available in restricted patient data from the VA’s Corporate Data Warehouse (CDW) database.

97 MATHEMATICS AND COMPUTING

Postearthquake Damage Mapping via Remote Sensing: Lessons From the 2023 Türkiye Disaster

This review addresses the urgent need for scalable, accurate, and reproducible remote sensing solutions following the February 2023 Türkiye earthquakes. It synthesizes the contributions of five peer-reviewed studies published in the IEEE JSTARS Special Issue on postearthquake damage and risk assessment. These studies cover areas such as damage classification with deep learning, fusion of multisource remote sensing data, creation of benchmark datasets, detailed damage mapping, and analysis of geophysical signals using outgoing longwave radiation. The article summarizes the methodological approaches and the practical relevance of the reviewed studies for detecting, evaluating, and quantifying damage, and outlines key challenges, including model generalization, class ambiguity, and data integration. It also discusses emerging trends, including explainable artificial intelligence, multimodal data fusion, and open-data platforms. This synthesis provides a foundation for building robust, interpretable, and real-time disaster response systems and aims to guide future research in earthquake-related Earth observation and rapid damage assessment.

Taskin, Gulsen [Istanbul Technical University] (OR

The need for standardization and improved open (meta)data practices in metaproteomics

Metaproteomics enables functional insight into microbial communities by identifying and quantifying proteins in complex samples. Yet, heterogeneous analytical workflows and the lack of standardization across experimental and bioinformatics stages hinder reproducibility and comparability, limiting integration with other omics data. We here present a community-developed reporting checklist tailored to the specific needs of metaproteomics. We also outline current efforts to enable structured and interoperable metadata capture, drawing on standards from proteomics and microbiome research wherever possible. By promoting transparent reporting and advancing metadata practices, our recommendations aim to align metaproteomics more closely with FAIR principles and support reproducible and interoperable research practices.

Armengaud, Jean [Universite Paris-Saclay, France]

Open-Source Data for MAC-POSTS: Mobility Data Analytics Center - Prediction, Optimization, and Simulation Toolkit for Transportation Systems

MAC-POSTS (Mobility Data Analytics Center - Prediction, Optimization, and Simulation toolkit for Transportation Systems) is a toolkit for dynamic transportation network modeling. Developed by the Mobility Data Analytics Center (MAC) at Carnegie Mellon University, this package implements many classic dynamic transportation network models, as well as new models proposed by MAC members. It has served as one building block for many other models and research projects. As such, this package used to be treated as an internal research project of the MAC lab, and admittedly, the code base is messy, and the interface is hard to use. However, we are working hard to make it a generally usable and useful toolkit for dynamic transportation network modeling. We would really appreciate any feedback, comments, suggestions, or criticisms.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

BOSC 2025, the 26th Bioinformatics Open Source Conference

The 26th annual Bioinformatics Open Source Conference (BOSC 2025, open-bio.org/events/bosc-2025) brought its community-driven focus on open-source bioinformatics and open science to the 2025 conference on Intelligent Systems for Molecular Biology and the European Conference on Computational Biology (ISMB/ECCB 2025). Since its launch in 2000, BOSC has been the premier annual meeting covering open-source bioinformatics and open science. Framed by two keynote addresses and a thought-provoking panel discussion, the two-day conference included sessions dedicated to open data, analytic tools and pipelines, workflow platforms, knowledge representation, and the application of AI/ML. The first keynote talk was delivered by Christine Orengo: “Working together to develop, promote and protect our data resources: Lessons learnt developing CATH and TED.” A joint session with the Bio-Ontologies and Knowledge Representation (BOKR) track the second day of BOSC started with a keynote talk by Chris Mungall entitled “Open Knowledge Bases in the Age of Generative AI”. A closing panel on Data Sustainability, moderated by Mónica Muñoz Torres, featured panelists Scott Edmunds, Varsha Khodiyar, Tony Burdett, Nicky Mulder, and Chris Mungall. This year, the CollaborationFest collaborative work event that typically precedes or follows ISMB was incorporated as part of the main conference and organized by BOSC with help from the Function and 3D-SIG tracks.

bioinformatics

Circularity Futures Workshop Series: Summary Report

The aim of this report is to synthesize key feedback received from the three-part Circularity Futures workshop series held in Spring 2024. The workshop series was conducted by the National Renewable Energy Laboratory (NREL) on behalf of U.S. Department of Energy, Office Energy Efficiency and Renewable Energy (EERE), and was broken into three workshops: Workshop 1 - Circularity Analysis Needs and Priorities; Workshop 2 - Circularity Metrics and Indicators; and Workshop 3 - Circularity Data. Together, the workshops focused on identifying the existing priorities and gaps in the circularity modeling space, understanding different stakeholders' use and interpretation of circularity metrics and indicators, identifying common data gaps and data quality challenges, and assessing the robustness of available solutions. The workshop series brought a diverse group of stakeholders - including representatives from U.S. government offices, national labs, nonprofit organizations, industry, and academia - to collect first-hand feedback on needs, priorities, challenges and opportunities in the circularity modeling and analysis space. The workshop discussions highlighted numerous common needs, priorities and challenges among the interviewed groups. Several topics were frequently discussed, including: 1) Circularity as a pathway for sustainable economic growth: While circularity is generally defined in terms of resource conservation and reducing wasteful disposal of materials, participants agreed that circular strategies should serve broader economic, environmental, and social goals. It is therefore crucial for circularity analysis to look beyond waste reduction and instead evaluate a variety of impact metrics such as cost savings, job creation, air quality, and pollutant emissions. Mutli-criteria decision-making frameworks may be useful for making sense of disparate metrics and evaluating tradeoffs between impact categories.; 2) Economic and social factors are not well understood: Underdevelopment of existing end-of-life (EOL) management infrastructure, inconsistent standardization codes and policy space in reusing recycled content, and suboptimal collection and sorting strategies collectively contribute to uncertainty about the economic potential of circular pathways. The latter observation is consistent among all technologies but more emphasized for renewable energy systems. Social impacts of circularity practices are less understood and less researched than other sustainability aspects.; 3) Inconsistent methods for assessing emerging technologies: LCA and TEA results vary widely depending on the assumptions made with regards to market adoption of new technologies. Emerging technologies suffer limited availability of data needed to conduct a robust circularity analysis. Yet, understanding projected impacts of proposed nascent technology is a key need for different stakeholder groups.; and 4) Lack of temporally and geospatially explicit data: There is a need for open data that represents variations in circularity technologies over time and location. The lack thereof leads to aggregated and potentially misrepresented results in circularity analysis. Sensitivity analyses should be included to verify whether options perceived as more sustainable align with real-world practices.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Data Driven Commercial Building Energy Code Compliance and Technology Inventory for New York City

Building Performance Standards (BPS) are gaining national traction. A BPS will require new processes in the design, construction, and operation of buildings that take the occupants into account and enable predictive analysis to ensure compliance with current and future GHG emissions caps. In New York City, most buildings over 25,000 square feet will be regulated by a BPS starting in 2024, regardless of whether it is new construction permitted under current energy codes or an existing building. This research is one of the first to begin the evaluation of a long-term series of building policies in the context of an open data ecosystem, in cooperation with city agencies. Existing building policies enacted in NYC have ranged from building energy benchmarking and labeling to energy audits to the regulation of GHG emission in buildings. Through the development of a dataset related to building technologies and energy consumption, this project can help to evaluate if meaningful conclusions can be drawn for the data that has been largely self-reported in compliance with city regulations. This project will also provide lessons learned from a deep dive into these types of datasets to provide best practices for municipalities or states seeking to embark on policies like those enacted in NYC. In addition, a Building Automation System (BAS) Stretch Standard of Care (SSOC) for owners, designers, and building operators will enable the measurement and predictive analysis of energy consumption and GHG emissions at the plant, system, or component level, in anticipation of regulated GHG limits on buildings based on energy use. The SSOC is expected to be suitable for use on a national level. The primary feature of an SSOC is a standardized format for a set of BAS points that can be used to control and to gather data from individual plants, systems, or components that are related to building energy consumption. This project examined how measurements compare to prescriptive or simulation-based energy code targets, finding little correlation between predictive 8760-hour energy modeling and actual energy consumption for a small sample (n=27) of buildings constructed after 2015. Other analysis found that, while large multifamily housing (MFH) buildings showed a general trend similar to predicted reductions in energy use from the implementation of model commercial energy codes, this trend was not evident in the office, K-12 school, and hotel use groups in NYC. No upward or downward trends in energy consumption were found when buildings were grouped by size. Energy audit data were analyzed and it appears that there is bias by audit company on measures recommended to clients. Further research should be performed to cross-analyze this with other attributes, such as building size, vintage, and number of stories. Analysis found that for 281 buildings that were permitted and completed after 2015 and had submitted benchmarking data in 2022, between 81% and 96% (by use group) were found to be in compliance with the 2024 to 2029 NYC BPS emission caps, and between 55% and 89% were in compliance with the 2030-2034 caps. This work is beneficial to the public in helping policymakers and building stakeholders better understand the wide-ranging implications of a BPS.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Ten questions concerning low-cost indoor air quality sensors: Perspectives from research and practice

Low-cost indoor air quality (IAQ) sensors are increasingly being used in homes and commercial and public buildings, driven by growing concerns about the impact of air on health, cognitive performance, and occupant wellbeing. These sensors offer a potentially transformative opportunity to increase spatial and temporal coverage of IAQ monitoring at a fraction of the cost of conventional reference instruments. However, their widespread use raises questions around accuracy, calibration, placement, data handling and interpretation, and integration into existing standards and workflows. This paper presents ten critical questions concerning the use of low-cost IAQ sensors in buildings, drawing on the latest empirical research, field deployments, and emerging practice. It discusses potential frameworks for deployment and evaluation, examines current sensor capabilities for measuring common pollutants, identifies methodological gaps in validation and uncertainty quantification, and outlines the extent to which existing IAQ standards can accommodate sensor-based evidence. The paper also explores how monitoring needs and deployment models vary by building type, the potential of real-time IAQ data to support building operations, and the ethical and legal implications of widespread sensor use. While significant challenges remain in ensuring data quality and building stakeholder trust, new applications are emerging through open data initiatives and advances in analytics and visualization. As the technology, science, and standards co-evolve, low-cost IAQ sensors are poised to become integral to routine building operation, building science, and environmental health research.

Parkinson, Thomas

Open Source Synergy: Developing and Validating PMU Data Analysis Techniques Using Open Source Tools and Datasets

This paper presents an exploration into the development and validation of data analysis approaches for Phasor Measurement Units (PMUs) using open-source datasets and tools. Various methods for event detection, event classification, frequency response, and oscillation analysis were tested. We leverage the capabilities of Archive Walker (AW), the Frequency Response Analysis Tool (FRAT), and the Oscillation Baselining and Analysis Tool (OBAT), all open-source tools, for efficient processing and analysis of synchrophasor data. The open-source Transmission Signature Library (TSL) dataset was employed as a dataset for a comprehensive evaluation to assess the performance and reliability of the proposed methods.

PMU, event analysis, oscillation, Frequency Respon

Visual Systems Mapping to Define and Compare Woody Biomass LCAs for Sustainable Systems

The challenge addressed in this research centres on the need to choose between several biomass sources and energy production processes, while supporting rural economies and resilience of forest systems. A key barrier to effective decision-making for strategies using biomass is the lack of standardized and transparent life cycle assessment (LCA) baselines. These baselines are critical for assessing the impacts of biomass strategies but often vary due to regional factors and chosen simplifying assumptions of the LCAs. However, omitting key variables can mean the LCA omits key feedback and balancing loops relevant to fully assessing impacts of the change or test scenario. To address these complexities, this project employs a systems engineering approach: visual systems mapping. This technique is used to define the boundaries and dynamic behaviours of LCA baselines, enhancing transparency. By examining five literature sources and their documented baseline scenarios, the systems mapping case-studies demonstrates an approach to documenting and archiving these baselines. Recommendations are that visual systems mapping should be used to document key assumptions, such as baselines, of LCAs. Further, where possible open data repositories should hold key information about LCA baselines and reproducible workflows (e.g., using open-source tools) should be used to improve transparency and comparability in LCAs. Given the consensus within the broader scientific community on the importance of replicable data practices, this research reinforces the need for standardized frameworks and systems engineering tools in LCAs. This research demonstrates a pathway to more transparent, standardized, and comparable LCAs, that may bolster decisions for biomass systems.

Davis, Maggie [ORNL] (ORCID:0000000181319328)

Towards a Public Event Display for DUNE

The Deep Underground Neutrino Experiment (DUNE) is a next generation long baseline neutrino experiment based at Fermilab, with a near detector near the beam target and a Far Detector (FD) in South Dakota. As the experiment prepares for its first data runs, creating pathways for public engagement and data transparency is essential. We present the first-ever DUNE event display designed for public outreach and education. Developed using data from the ProtoDUNE detectors at the CERN Neutrino Platform, this tool provides an intuitive and interactive interface that allows non-experts to visualise and explore particle interactions in a Liquid Argon Time Projection Chamber (LArTPC). By translating raw experimental data into a browser-accessible format, we establish the essential infrastructure for DUNE’s pathway to open data. This talk will detail the technical development of the display, its current implementation with ProtoDUNE data, and the strategic roadmap for integrating it into DUNE’s long-term open-access framework.

Sabater, Eva [U. Sussex (main)] (ORCID:00090001748

Accelerated data-driven materials science with the Materials Project

The Materials Project was launched formally in 2011 to drive materials discovery forwards through high-throughput computation and open data. More than a decade later, the Materials Project has become an indispensable tool used by more than 600,000 materials researchers around the world. This Perspective describes how the Materials Project, as a data platform and a software ecosystem, has helped to shape research in data-driven materials science. We cover how sustainable software and computational methods have accelerated materials design while becoming more open source and collaborative in nature. Next, we present cases where the Materials Project was used to understand and discover functional materials. We then describe our efforts to meet the needs of an expanding user base, through technical infrastructure updates ranging from data architecture and cloud resources to interactive web applications. Finally, we discuss opportunities to better aid the research community, with the vision that more accessible and easy-to-understand materials data will result in democratized materials knowledge and an increasingly collaborative community.

Horton, Matthew K