Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “FAIR data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

DOE FAIR Surrogate Benchmarks Supporting AI and Simulation Research (SBI Surrogate Benchmark Initiative) (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia(UVA). SBI repositories include data, code, and all relevant collateral artifacts, that the science and engineering community needs to use and reuse these data sets and surrogates. SBI repositories generate active research from both participants in SBI and the broader AI and domain science communities. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and capture them as surrogate benchmarks with a rich set of metadata, covering. Data; Model; Metrics specification; Machine specification; Science, Speed, Power Results, We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non Surrogate benchmarks that have many common features and similar issues regarding FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, benchmarks have datasets, models, and metadata, and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates, including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

FAIR Surrogate Benchmarks Supporting AI and Simulation Research (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia (UVA). SBI repositories include data, code, and all relevant collateral artifacts that the science and engineering community need to use and reuse these data sets and surrogates. SBI repositories generate active research from both the participants in SBI and the broad community of AI and domain scientists. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and captures them as surrogate benchmarks with a rich set of metadata covering: Data; Model; Metrics specification; Machine specification; and Science, Speed, and Power Results. We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non-Surrogate benchmarks that have many common features and similar issues as regards FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, Benchmarks have datasets, models, and metadata and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

Alaska's Rural Building Stock: a Validation Study Using ResStock and Field Data

The availability of accurate national data on demographics, building stock, and energy use is vital for modeling residential buildings and evaluating decarbonization strategies. However, rural and Indigenous populations, including those in rural Alaska, are typically underrepresented in these datasets. These communities face unique challenges due to their remote locations, severe weather conditions, and limited access to resources, resulting in high energy burden. This report examines how rural Alaskan communities are underrepresented in the ResStock housing model and highlights the need for improved data to address their unique housing and energy challenges. Thus, this report examines the representation of rural Alaskan communities within the national housing stock model, ResStock. A validation study was conducted, considering ResStock, Field Data and Aerial and 3D-view data collection (A3DDC) datasets. The validation process started by using the down selecting approach on the ResStock building stock dataset. For the purpose of this study, only the rural Alaska Boroughs and Census areas located in ASHRAE IECC Climate Zone 8 were considered to ensure a more accurate and fair comparison with the field data, which was collected in rural areas located in climate zone 8, specifically within the Nome Census area. While ResStock may accurately represent several characteristics of the building stock for rural Alaska, some differences between modeled, field data, and aerial and 3D-view data collection datasets were identified. The following building characteristics have a high impact on modeled energy consumption and demonstrated large differences: Revisit heating setpoints and consider a substantially higher setpoint distribution, it could potentially address "missing loads" if this is the case. Develop and include Toyo heating in future modeling for ResStock and EnergyPlus. Remove natural gas as a water heater fuel type outside of North Slope County. Foundation type updated to have more crawlspaces rather than basements. Infiltration rates need reexamination for a larger distribution toward higher infiltration rates. Include more vinyl and less brick in exterior wall type and revise wall color for greater proportion of light rather than dark color. Roof material revised from majority shingles to majority metal. Update number of occupants to higher number of occupant count. Building orientation represents a higher proportion of south facing buildings rather than relatively equal. The findings suggest that updating ResStock's probability logic could better represent rural Alaskan buildings. ResStock can be utilized to identify the best upgrades or energy efficiency and energy efficiency improvements, helping community leaders in making more informed decisions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Confronting Large‐Eddy Simulations With Stereo Camera Data by Means of Reconstructed Hemispheric Cloud Size Distributions

High-resolution hemispheric camera images at a meteorological site in western Germany are used to analyze the multi-dimensional spatial characteristics of continental cumulus cloud fields, and to evaluate Large-Eddy Simulations on this aspect. Traditional non-hemispheric cloud-detecting instruments provide additional reference data. The main model-observation comparison focuses on cloud size distributions (CSDs), employing two methods: (a) directly using three-dimensional model fields, direct CSDs, and (b) using rendered hemispheric images of the model fields as produced by a camera simulator based on path-tracing. In the latter method, both the real and rendered images are used to three-dimensionally reconstruct the cloud fields, yielding hemispheric CSDs. Advantages of hemispheric comparisons over more classic approaches include (a) fair comparisons between model and data, and (b) full use of the enhanced resolutions and hemispheric spatial coverage of the camera imagery. Basic evaluation of the simulations demonstrates good agreement on thermodynamic structure and its diurnal cycle. Cloud heights and cloud cover are intercompared between the model, camera data and other instrumentation, providing insight into their structural differences. A consistent alignment is found between the hemispheric CSDs from both the model and the cameras. Power law fits reveal structurally lower exponents in hemispheric CSDs compared to non-hemispheric CSDs, which particularly caution against directly comparing hemispheric CSDs to non-hemispheric distributions. This result is robust for sample size and fitting method. These findings inform future use of hemispheric camera systems for studying cumulus cloud field morphology and model evaluation.

54 ENVIRONMENTAL SCIENCES↗

FAIR Framework for Physics-Inspired AI in High Energy Physics (Final Technical Report)

The main deliverable of this proposal was to publish data from high energy physics experiments in a FAIR format so that non-specialists could develop machine learning technologies using our data. The Minnesota team of Profs. Cushman, Furmanski and Rusack, from the high energy experiments CDMS, Micro-Boone and CMS, respectively, and Prof J. Sun from Computer Science worked to organize the data, to provide code to access the data, and where relevant provide documentation describing the data. The FAIR4HEP collaboration was formed with groups from UC San Diego, MIT, and the University of Illinois, with the principal investigator was Dr. Huerta. Collectively we collaborated on the publication of datasets from the LHC experiments. Members of the Minnesota group contributed to the common papers published by the collaboration

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Indications for Freeze-Out of Charge Fluctuations in the Quark-Gluon Plasma at the LHC

The D-measure of net-charge fluctuations quantifies the variance of net charge in strongly interacting matter. It was introduced over 20 years ago as a potential signal of quark-gluon plasma (QGP) in heavy-ion collisions, where it is expected to be suppressed due to the fractional electric charges of quarks. Measurements have been performed at RHIC and LHC, but the conclusion has been elusive in the absence of quantitative calculations for both scenarios. We address this issue by employing a recently developed formalism of density correlations and incorporate resonance decays, local charge conservation, and experimental kinematic cuts. We find that the hadron gas scenario is in fair agreement with the ALICE data for $\sqrt{{𝑠}_{\textrm{NN}}}$ = 2.76 TeV Pb–Pb collisions only when a very short rapidity range of local charge conservation is enforced, while the QGP scenario is in excellent agreement with experimental data and largely insensitive to the range of local charge conservation. A Bayesian analysis of the data utilizing different priors yields moderate evidence for the freeze-out of charge fluctuations in the QGP phase relative to hadron gas. The upcoming high-fidelity measurements from LHC Run 2 will serve as a precision test of the two scenarios.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Challenges and Vision for Standardization of Biopolymer Data Sets for Machine Learning

Machine learning (ML) is transforming materials research, yet potential for biopolymer discovery remains constrained by fragmented data and nonstandardized reporting. Biopolymers differ significantly from synthetic polymers, requiring specialized approaches to represent their biosynthetic origins, hierarchical structures, and application-specific metrics. In this Perspective, we identify three core challenges limiting biopolymer representation: information encoding, data quality, and data sharing. We describe the most pressing issues and propose commensurate approaches to address each key challenge. Recommendations include the design and adoption of biopolymer-specific fingerprinting and representation frameworks, development of hybrid human-large language model (LLM) data extraction strategies, and expanding Findable, Accessible, Interoperable, Reusable (FAIR)-compliant repositories. We propose a robust foundation to define interoperable, high-quality data sets that capture the full context of biopolymer materials. Standardized metadata, shared ontologies, and community-driven infrastructure would enable scalable, reproducible workflows and accelerate the ML-driven development of biopolymers.

36 MATERIALS SCIENCE↗

PDB-IHM: A System for Deposition, Curation, Validation, and Dissemination of Integrative Structures

Structures of many large biomolecular assemblies are now being determined using integrative approaches. In these approaches, information derived from multiple experimental and computational methods is combined to compute three-dimensional structures of multi-protein complexes and other macromolecular machines. A standalone prototype data resource for integrative structures called PDB-Dev was built, based on recommendations of the Integrative and Hybrid Methods (IHM) Task Force of the Worldwide Protein Data Bank (wwPDB). This effort included developing data standards and software tools for collecting, curating, validating, visualizing, archiving, and disseminating integrative structures that span diverse spatiotemporal scales and conformational states. Mechanisms have been created to validate integrative structures based on the experimental data underpinning them. Building upon this foundational framework, PDB-Dev has been further expanded to handle large dynamic macromolecular systems and integrative structures that combine, for example, experimental restraints with atomic coordinates computed by machine learning algorithms. Data standards and supporting tools have also been extended to capture information about biomolecular dynamics, such as conformational transitions and related kinetic data derived from biophysical methods. Recently, PDB-Dev was unified with the PDB archive and rebranded as PDB-IHM (pdb-ihm.org), further promoting FAIR (Findable, Accessible, Interoperable, and Reusable) principles of data stewardship for integrative structural biology.

IHMCIF↗

Workflow Provenance in the Computing Continuum for Responsible, Trustworthy, and Energy-Efficient AI

As Artificial Intelligence (AI) becomes more pervasive in our society, it is crucial to develop, deploy, and assess Responsible and Trustworthy AI (RTAI) models, i.e., those that consider not only accuracy but also other aspects, such as explainability, fairness, and energy efficiency. Workflow provenance data have historically enabled critical capabilities towards RTAI. Provenance data derivation paths contribute to responsible workflows through transparency in tracking artifacts and resource consumption. Provenance data are well-known for their trustworthiness helping explainability, reproducibility, and accountability. However, there are complex challenges to achieve RTAI, which are further complicated by the heterogeneous infrastructure in the computing continuum (Edge-Cloud-HPC) used to develop and deploy models. As a result, a significant research and development gap remains between workflow provenance data management and RTAI. In this paper, we present a vision of the pivotal role of workflow provenance in supporting RTAI and discuss related challenges. We present a schematic view between RTAI and provenance, and highlight open research directions.

Santos Souza, Renan↗

Supporting ARPA-E Power Grid Optimization (Final Report)

Pacific Northwest National Laboratory (PNNL), Arizona State University (ASU), Georgia Institute of Technology (Georgia Tech), Los Alamos National Laboratory (LANL), National Renewable Energy Laboratory (NREL), Texas A&M University (TAMU), The University of Texas at Austin (UT), and the University of Wisconsin-Madison (UW-M) supported the ARPA-E Grid Optimization (GO) Competition by providing a common problem formulation, data format, datasets, evaluation mechanism, scoring, rules, and results that resulted in the awarding of $\$9.24$ million dollars to teams from academia, industry, and national labs for solving three sets of increasingly difficult non-linear, security- constrained AC Optimal Powerflow (AC-OPF) optimization problems in order to increase the efficiency of the US Electric Grid. It is estimated that a 1% increase in efficiency can save $\$1$ billion. Current industry practices typically use a linear DC model (DC-OPF) in order solve the OPF problem within the time constraints of the operation schedule. The GO Competition challenges the best power engineers, mathematicians, and computer scientists to make possible operational decisions based on accurate physical models. To accomplish this, the GO Competition created a series of Challenges and funded teams to produce the best solver. Challenge 1 was to solve the security constrained Alternating Current Optimal Power Flow (ACOPF) problem. Challenge 2 extended that to by adding adjustable transformer tap ratios, phase shifting transformers, switchable shunts, price-responsive demand, ramp rate constrained generators and loads, and fast-start unit commitment (UC). Furthermore, Challenge 2 was a maximization problem while Challenge 1 was a minimization problem. While Challenge 3 was being developed, the entrants were invited to find better solutions to the Challenge 2 synthetic datasets with no restrictions on time, hardware, or algorithms. The Challenge 2 solutions turned out to be very good. Challenge 3 expanded the Challenge 2 problem further by using multiperiod dynamic markets, including advisory models for extreme weather events, day-ahead markets, and the real-time markets with an extended look-ahead. These problems included active bid-in demand and topology optimization. Together the Challenges used nearly 30 million CPU hours. Since each team was working on the same problem, using the same data, and running on the same hardware, fair comparisons could be drawn as to the best solver. The datasets were varied enough, however, that the best solver for one dataset was not necessarily the best at another, so cumulative scores were used. The process was managed by the PNNL maintained website https://GOCompetition.energy.gov, where Entrants could find information about the problem, the data, the rules, submit their solver for evaluation, and see the scores of all the competing teams on a Leaderboard. Interest was world-wide but only American teams were eligible for prizes. The Competition has produced 34 journal articles 115 papers and been cited over 500 times in the literature, including 12 dissertations (4 from foreign countries; Columbia (2), Germany, and Italy) and 3 from the DOE ExaScale project. Software developed by Pearl Street Technologies for Challenges 1 and 2 is now deployed by Southwest Power Pool (SPP) and Midcontinent Independent Service Operator (MISO). Other teams have received inquiries from venture capitalists. Google DeepMind has thanked the Competition for making the datasets developed for the Competition public. They are using it to train machine learning models. The larger datasets have billions of unknowns to be solved for, but only a small percent matter in the final solution. Knowing what unknowns are important can dramatically speedup the solution.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Exploring the role of clouds in offshore wind potential off the US West Coast in a changing climate

To meet US goals of deploying additional wind energy as part of the decarbonization strategy, wind plants are being planned for the deep water offshore the western US. The wind flow in that region is complex due to the proximity to the coast, cold water upwelling, and persistent stratiform clouds that interact with radiation in ways that have the potential to destabilize the atmosphere. That flow has the potential to change with a changing climate. To address these issues, we assess the flow and the clouds in that region using downscaled climate model data, under both historic climate (1975–2005) and projected future (2025–2055) conditions. We note that the climate simulations agree fairly well with the cloud patterns observed by satellite data in the nearshore and offshore regions. We then assess the projected changes in clouds, wind speed, and other important variables, noting that our simulations project that the predominant north/northwesterly low-level jet is expected to strengthen and clouds are likely to be commensurately enhanced, although projected changes are within about 10% of current conditions. Our examination of the dynamics associated with the changes in the climate simulations provides confidence in the dynamical consistency of these projected changes.

17 WIND ENERGY↗

Analysis of static Wilson line correlators from lattice QCD at finite temperature with T -matrix approach

The thermodynamic T-matrix approach is used to study Wilson line correlators (WLCs) for a static quark-antiquark pair in the quark-gluon plasma (QGP). Selfconsistent results that incorporate constraints from the QGP equation of state can approximately reproduce WLCs computed in 2+1-flavor lattice-QCD (lQCD), provided the input potential exhibits less screening than in previous studies. Utilizing the updated potential to calculate pertinent heavylight T-matrices we evaluate thermal relaxation rates of heavy quarks in the QGP. We find a more pronounced temperature dependence for low-momentum quarks than in our previous results (with larger screening), which turns into a weaker temperature dependence of the (temperature-scaled) spatial diffusion coefficient, in fair agreement with the most recent lQCD data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Cross sections for the formation of Rb84m,g, Rb83, and Rb82m in Sr86(d,x) reactions up to deuteron energies of 49 MeV: Competition between α-particle and multinucleon emission processes

Cross sections of Sr86(d,x) reactions leading to the products Rb84m,g, Rb83, and Rb82m were measured by the stacked-sample activation technique up to deuteron energies of 49 MeV. Nuclear model calculations were performed using the codes talys and empire, which combine the statistical, precompound, and direct interaction components. In all cases, the empire results were much higher than the talys calculation. Fairly good agreement was obtained between measured data and the talys calculation after some optimization of the input model parameters. Insight into competition between α-particle and multinucleon emission in the Y88 compound-nucleus system was also gained.

59 ≤ A ≤ 89↗

rcsb-api : Python Toolkit for Streamlining Access to RCSB Protein Data Bank APIs

The Protein Data Bank (PDB) was founded in 1971 as the first open-access digital data resource in biology to serve as the single global archive for three-dimensional (3D) macromolecular structure data. Current PDB holdings exceed 230,000 experimentally determined structures of proteins, nucleic acids, viruses, and macromolecular machines. The RCSB Protein Data Bank RCSB.org research-focused web portal facilitates search, analyses, and visualization of every PDB structure along with more than one million Computed Structure Models from AlphaFold DB and the ModelArchive. It is powered by a set of publicly available Application Programming Interfaces (APIs) that both support RCSB.org users and provide programmatic access to PDB data. Given the breadth and levels of granularity encompassed in this rich data collection, efficiently accessing the information programmatically may be challenging for new users. RCSB PDB has developed a Python software package, rcsb-api , that facilitates easy and efficient use of RCSB PDB APIs within a Python environment. This software tool is designed to streamline access to the extensive corpus of data housed within the PDB, enabling researchers to search, retrieve, and analyze 3D biostructure data seamlessly. Its use will accelerate research in structural biology, molecular biology and biochemistry, drug discovery, and bioinformatics by providing more efficient tools for data integration and analysis. The new toolkit is available on GitHub (github.com/rcsb/py-rcsb-api) and published to the public Python package repository (PyPI) to foster wider usage and support basic and applied research in fundamental biology, biomedicine, and the energy sciences.

FAIR principles↗

Electricity Rate Designs for Large Loads: Evolving Practices and Opportunities

Electricity demand from large load customers such as data centers is projected to grow significantly in the near term. While data centers play an important role in advancing technology innovation and economic growth in the United States, data center energy needs present challenges and opportunities for electricity supply and infrastructure. This technical brief serves as a foundation for the discussion of issues and sharing of perspectives among utilities, regulators, large load customers, and other stakeholders. As utilities and regulators explore rate structures to address growing data center electricity demand, several issues have emerged: -Fair allocation of electricity system costs to large-load customers without unfair shifting of costs to other customers -Appropriate mitigation of the financial risks associated with stranded assets from underutilized utility system investments -Mitigation of operational and resource adequacy risks if electricity demand exceeds supply -Appropriate risk-sharing in commercializing newer electricity technologies such as advanced geothermal, small modular reactors, and long duration energy storage -Accommodating the diverse needs of large-load customers, such as having the option to match electricity consumption with output from carbon-free resources or using onsite generation to provide system capacity The technical brief also identifies key design elements that aim to address these issues and uses leading examples from pending and approved rate structures, agreements, and special contracts to ground the elements in practice.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Toward equitable environmental exposure modeling through convergence of data, open, and citizen sciences: an example of air pollution exposure modeling amidst increasing wildfire smoke

Exposure modeling is critical in environmental epidemiology and human health but may face challenges (e.g., skewed data, unequal error, context-insensitive validation, and computational demands). Modeling decisions reflect the intended use of the models and the values that modelers prioritize. We aimed to provide a conceptual framework and machine learning (ML) modeling protocols that address these issues. With 500m-gridded hourly PM 2.5 and O 3 levels in Illinois before, during, and after the 2023 Canadian wildfire season as a motivating example, we conducted modeling experiments to evaluate modeling methods, guided by three domains we propose based on theories of science: 1) Data Diversity, leveraging open and citizen science data to enhance inclusivity, parsimony, and representativeness; 2) Equitable Accuracy, ensuring fairly distributed uncertainties across subpopulations; and 3) Sustainable Modeling, balancing accuracy with reducing computational demands to promote accessibility for under-resourced researchers. Here, we found that ML with publicly available data can achieve high accuracy. Depending on methods, performance may vary substantially, even with identical input data. Large but skewed data may reduce performance. Misuse of cross-validation protocols can underestimate prediction error; although we observed R 2 s of ∼98 %, the modeled estimates varied significantly, indicating the need for careful model validation. By using new modeling protocols including representativeness-considered training and validation data and a new loss function, we achieved high agreement between estimates and ground-based measurements (e.g., R 2 = ∼90 % for PM 2.5 ; ∼80 % for O 3 ), equally distributed errors across sociodemographic strata and urban–rural divides, and reduction in computation time—from several weeks or months to a few days.

Exposure assessment↗

An AI-Enabled Chat Bot for DuraMAT

The DuraMAT Data Hub has evolved to meet the ever expanding research demands by integrating a chatbot interface that enhances FAIR compliance and simplifies discovery across our growing archive of projects and datasets. With the Data Hubs increasing use by both consortium researchers and the international community, traditional navigation methods have become unwieldy. Building on last year's feasibility study, we now report the successful implementation of the chatbot system, detailing its architecture and demonstrating its ability to deliver a more intuitive and rewarding user experience.

14 SOLAR ENERGY↗

The DECADE cosmic shear project III: validation of analysis pipeline using spatially inhomogeneous data

We present the pipeline for the cosmic shear analysis of the Dark Energy Camera All Data Everywhere (DECADE) weak lensing dataset: a catalog consisting of 107 million galaxies observed by the Dark Energy Camera (DECam) in the northern Galactic cap. The catalog derives from a large number of disparate observing programs and is therefore more inhomogeneous across the sky compared to existing lensing surveys. First, we use simulated data-vectors to show the sensitivity of our constraints to different analysis choices in our inference pipeline, including sensitivity to residual systematics. Next we use simulations to validate our covariance modeling for inhomogeneous datasets. Finally, we show that our choices in the end-to-end cosmic shear pipeline are robust against inhomogeneities in the survey, by extracting relative shifts in the cosmology constraints across different subsets of the footprint/catalog and showing they are all consistent within 1σ to 2σ. This is done for forty-six subsets of the data and is carried out in a fully consistent manner: for each subset of the data, we re-derive the photometric redshift estimates, shear calibrations, survey transfer functions, the data vector, measurement covariance, and finally, the cosmological constraints. Our results show that existing analysis methods for weak lensing cosmology can be fairly resilient towards inhomogeneous datasets. This also motivates exploring a wider range of image data for pursuing such cosmological constraints.

79 ASTRONOMY AND ASTROPHYSICS↗