Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data systems standards”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Physics with high-luminosity proton-nucleus collisions at the LHC

The physics case for the operation of high-luminosity proton-nucleus (pA) collisions at the CERN LHC is reviewed. The collection of $\mathcal{O}$(1–10 pb −1 ) of proton-lead (pPb) collisions at the LHC will provide unique physics opportunities in a broad range of topics including proton and nuclear parton distribution functions (PDFs and nPDFs), generalised parton distributions (GPDs), transverse momentum dependent PDFs (TMDs), low-x quantum chromodynamics and parton saturation, hadron spectroscopy, baseline studies for quark-gluon plasma and parton collectivity, double and triple parton scatterings, photon–photon collisions, and physics beyond the Standard Model; which are not otherwise as clearly accessible by exploiting data from any other colliding system at the LHC. This report summarises the accelerator aspects of high-luminosity pA operation at the LHC, as well as each of the physics topics outlined above, including the relevant experimental measurements that motivate much larger pA datasets than collected to date.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Ectomycorrhizal effects on decomposition are highly dependent on fungal traits, climate, and litter properties: A model-based assessment. Dataset.

To simulate the effects of mycorrhizal fungi on soil organic matter cycling, we incorporated mycorrhizal processes into the Carbon, Organisms, Rhizosphere, and Protection in the Soil Environment (CORPSE) model to develop a new soil model Myco-CORPSE. The new model was calibrated and evaluated against soil measurements taken at temperate forests in New Hampshire (NH) and Georgia (GA). A series of scenario analysis were also conducted to explore the conditions under which ectomycorrhizal (ECM) N acquisition processes can induce different soil C accumulation in ECM systems compared to arbuscular (AM) systems.In this data package, we included:-The Python codes of the standard Myco-CORPSE model we developed: "Standard Myco_CORPSE python codes.zip". The main program is the "gradient_sim.py" which calculates the bulk soil microbes and CN content along a user defined gradient of clay, soil temperature, soil moisture and mycorrhizal dominance, and relies on two subprograms "CORPSE_deriv.py" and "CORPSE_integrate.py". "CORPSE_deriv.py" calculated the changes in all simulated soil stock within every time step and "CORPSE_integrate.py" integrate the changes in all simulated soil stock within simulated time period. The program "Plot.py" is used to plot the major outputs produced by the main program "gradient_sim.py".-The modified Python codes of Myco-CORPSE models with site-level environmental inputs (in NH and GA) used to conduct simulations in NH and GA sites: "NH_GA model simulations.zip". -The Python codes used to evaluate the Myco-CORPSE simulation outputs in NH and GA sites against site-level measurements: "Plot NH_GA simulation against measurements.zip". It includes both the evaluation Python code, the model outputs on NH and GA sites, and the measured soil properties in both sites.-The modified Python codes of Myco-CORPSE models "Scenario analysis_model simulations.zip" that is used to conduct scenario analysis of how different litter properties, mycorrhizal fungal traits, climate, and seasonal variation in temperature and vegetation phenology impact the mycorrhizal effects on soil CN properties. The sub file folder "Scenario analysis_litter traits" contains the codes for scenario analysis of different litter properties; The sub file folder "Scenario analysis_ECM types" contains the codes for scenario analysis of different ECM fungal traits; The sub file folder "Scenario analysis_climate&seasonality" contains the codes for scenario analysis of different climate and seasonalities;-"Scenario analysis_model results and plotting codes.zip" contains all the output files from the the scenario analysis of Myco-CORPSE model as described above and the plotting codes used to the generate the heatmaps and scatterplots shown in the manuscript "Ectomycorrhizal effects on decomposition are highly dependent on fungal traits, climate, and litter properties: A model-based assessment"The majority of the model outputs did not have specific geographic information or temporal coverage because the analysis we conducted are mainly hypothetical model simulations. We only provided geographic description, coordinates and temporal coverage for those soil measurements which we used for model evaluations (included in the "Plot NH_GA simulation against measurements.zip").

54 ENVIRONMENTAL SCIENCES↗

ATLAS-MAP: An Automated Test Station for Gated Electronic Transport Measurements

The diversification of electronic materials in devices provides a strong incentive for methods to rapidly correlate device performance with fabrication decisions. In this work, we present a low-cost automated test station for gated electronic transport measurements of field-effect transistors. Utilizing open-source PyMeasure libraries for transparent instrument control, the “ATLAS-MAP” system serves as a customizable interface between sourcemeters and samples under test and is programmed to conduct transfer curve and van der Pauw methods with static and sweeping gate voltages. Zinc oxide transistors of variable thickness (5, 10, and 20 nm) and channel size (50 μm to 3 mm, of equal length and width) were fabricated to validate the design. Standardization of testing procedures and raw data formatting enabled automated data analysis. A detailed list of parts and code files for the system are provided.

36 MATERIALS SCIENCE↗

Hardware-in-the-loop Laboratory Performance Verification of Flexible Building Equipment in a Typical Commercial Building

This project aims to develop high-resolution equipment performance and occupant data that quantifies demand flexibility in typical commercial buildings. The dataset documented in this report includes comprehensive time-series measurements from hardware-in-the-loop (HIL) experiments conducted across multiple testbeds designed to simulate realistic operational environments for HVAC systems. Specifically, it captures minute-by-minute high-resolution data on the performance of various typical HVAC systems, including a variable-air-volume (VAV) air handling unit (AHU) system with chillers and an ice tank in the Intelligent Building Agents Laboratory (IBAL) at the National Institute of Standards and Technology (NIST), a two-stage air-source heat pump (ASHP) at the NIST, and a water-source heat pump (WSHP) at Texas A&M University (TAMU). The data were generated under a range of controlled conditions reflecting different grid scenarios and climatic influences, as well as various control strategies, building types, occupancy patterns, and occupant behaviors. This dataset provides DE-EE0009153 Final Report 5 detailed insights into the demand flexibility of these systems, including their response to grid signals, occupant behaviors, energy consumption patterns, and operational efficiency under different conditions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Similarity Metric for Data Optimization and Efficient Training of Reactive Machine Learning Force Fields for Hydrocarbon Radiolysis

Radiolysis is a common approach to sterilize polymers, chemically modify them for upcycling, and accelerate their decomposition for recycling purposes. Reactive molecular dynamics (MD) simulations provide a powerful tool to generate atomic-level trajectories of the reactive processes and quantify radiolytic chemical degradation pathways. For this, machine learning (ML) surrogate models for reactive force fields with quantum mechanical accuracy are now widely used, which require ML training data sets that can provide information on atomic environments for target chemical systems. However, radiolysis chemistry can be highly complex and diverse, which poses significant challenges for generating training data to parametrize ML models. In this regard, we developed a method for optimizing the training data set using a cosine similarity metric to help guide training set selection for radiolysis of polyethylene, a model hydrocarbon polymer, as well as to enhance the transferability of our reactive ML force field (MLFF) to a variety of molecular and polymeric systems. Our approach performs atom-by-atom comparisons between local atomic environments to pinpoint important data points associated with rare and localized events, such as radiolysis damage within structures. We apply this approach to train the Chebyshev Interaction Model for Efficient Simulation (ChIMES) MLFF model, which expresses the atomic interaction potentials in terms of linear combinations of many-body Chebyshev polynomials. We first show that our method can reduce our training set size by ∼70% while improving overall accuracy compared to more standard MD model fitting approaches. We then validate our optimum model against diverse hydrocarbon simulation data, including simple alkanes and systems with unsaturated carbon bonds, over a wide range of thermodynamic conditions. Finally, we use our ChIMES model to perform MD simulations of radiolytic damage with large-scale systems that help avoid system size effects. Overall, our approach yields an MD force field that retains most of the accuracy of the underlying quantum method while yielding many orders of improvement in computational efficiency. In conclusion, our efforts will have impact on future hydrocarbon polymer radiolysis studies, where the chemical details of the polymer–radiation interactions can have a strong effect on the resulting products observed in experiments.

Hydrocarbons↗

Evaluating the factors influencing accuracy, interpretability, and reproducibility in the use of machine learning classifiers in biology to enable standardization

The complexity and variability of biological data has promoted the increased use of machine learning methods to understand processes and predict outcomes. These same features complicate reliable, reproducible, interpretable, and responsible use of such methods, resulting in questionable relevance of the derived. outcomes. Here we systematically explore challenges associated with applying machine learning to predict and understand biological processes using a well- characterized in vitro experimental system. We evaluated factors that vary while applying machine learning classifers: (1) type of biochemical signature (transcripts vs. proteins), (2) data curation methods (pre- and post-processing), and (3) choice of machine learning classifier. Using accuracy, generalizability, interpretability, and reproducibility as metrics, we found that the above factors significantly mod- ulate outcomes even within a simple model system. Our results caution against the unregulated use of machine learning methods in the biological sciences, and strongly advocate the need for data standards and validation tool-kits for such studies.

59 BASIC BIOLOGICAL SCIENCES↗

Evaluation and improvement of the parameterization of aerosol hygroscopicity in global climate models using in-situ surface measurements (Final Report)

Aerosols are tiny particles suspended on the atmosphere that can interact with incoming solar radiation and affect the Earth radiative budget. They do so by scattering and absorbing solar radiation, and these properties vary depending on the aerosol size and chemical composition. Moreover, by taking up water from the surrounding air, hygroscopic aerosol particles will grow in size and change their chemical composition, thus modifying their optical properties (scattering and absorption) and their final impact on radiative forcing calculations. An accurate knowledge of aerosol hygroscopicity is crucial for estimating the net radiative impact of aerosols. We took a three-pronged approach to improve our understanding of aerosol hygroscopicity and how it is implemented in Earth system models. In the first part of our project, we developed a benchmark dataset from existing aerosol hygroscopic growth measurements made by tandem nephelometer humidogram systems. We analyzed, using a standardized methodology, observations from 26 in-situ stations around the globe. Measurement data was collected from multiple data providers, reviewed and harmonized to create a consistent dataset of the scattering enhancement factor due to aerosol water uptake. This dataset is archived in several publicly available databases for use by interested researchers. In the second part of the project, we used the benchmark hygroscopicity dataset to perform a global study on aerosol hygroscopicity and aerosol optical properties. Measurements show a global picture of scattering enhancement with larger values for Arctic and marine sites and lower for urban and desert sites. We assessed the RH dependence of aerosol radiative forcing and showed that the overall effect of aerosol hygroscopicity on DARF is an increase in the absolute forcing effect (negative sign) by a factor of up to 4 compared to dry conditions (RH<40%). Finally, we explored using aerosol single scattering albedo (SSA) and scattering Angstrom exponent (SAE) as possible proxies for aerosol hygroscopicity. SSA showed more promise as a surrogate for the scattering enhancement factor than SAE, but neither was ideal. In the third part of the project we evaluated the output of ten Earth system models (ESMs) against the benchmark hygroscopicity dataset. ESMs utilize various schemes for aerosol hygroscopicity which had not been previously tested against observations on a global scale. We found that ESMs currently overestimate scattering enhancement due to hygroscopic growth. Model parameterizations of hygroscopicity and model chemistry are two main factors driving the observed diversity in hygroscopicity simulations among the models. In addition, our study makes several suggestions for modelers, including improving the parameterizations of organic and sea salt aerosol hygroscopicity. Future hygroscopicity model evaluation experiments should include the model data related to particles size which was not available for our study.

54 ENVIRONMENTAL SCIENCES↗

DOE Repository Metadata Profile (DRMP): A Metadata Framework for Advancing Interoperability and AI Readiness Across Scientific Repositories

The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Reducing Uncertainty of Fielded Photovoltaic Performance (Final Technical Report)

Improved analysis and reporting of photovoltaic (PV) field performance increases the certainty of owners and financiers that systems will perform as expected. Advanced module technologies (e.g., PERC, HJT, and bifacial) introduce new degradation mechanisms and performance characteristics. The FY19-21 Reducing Uncertainty project leveraged data from the ever-increasing PV fleet to develop models and understanding of the field performance of existing and new technologies. Specifically, we accomplished: report on field performance and degradation rates for high-efficiency silicon (HJT, PERC, IBC) and more conventional technologies; developed automated analysis techniques to quantify system performance (performance ratio, energy yield) and production shortfalls (soiling, degradation, availability); refined the RdTools software toolkit to bring standard, validated analysis techniques to bear on third-party data; analyzed and reported on large datasets including Treasury data and Lawrence Berkeley National Laboratory's Utility-Scale dataset to expand the high-quality degradation-rate histogram published previously; worked with industry partners and the DuraMAT data hub to enable private parties to share and aggregate PV production data anonymously, leveraging cloud-based data analysis infrastructure and publishing on US fleet-scale performance comprising over 7GW of operating systems. (https://www.nrel.gov/pv/fleet-performance-data-initiative.html). Through our industry collaborations we have engaged in NDA-covered data transfer with twelve PV fleet owners as of January 2022, with more agreements in negotiation. Our scalable cloud-based time series database contains over 30 billion rows (20TB) of PV time series data, representing over 1700 commercial and utility-scale systems, and over 7.2 GW of DC capacity (Fig 1). Initial field performance results have been distributed in several public reports. Because our fleet composition and data quality methods are continually improving, annual updates to these results are published to our PV Fleet webpage [ https://www.nrel.gov/pv/fleet-performance-data-initiative.html ] and DuraMAT data hub [DOI: 10.21948/1842958]. Another existing dissemination channel used for observed soiling losses is a map we maintain for soiling losses. Additional products developed include a report detailing fleet-wide performance index, availability, startup loss and snow loss factors, a detailed report on the 1603 grant dataset comprising over 100,000 PV systems with failure and performance details and a utility-scale report coauthored with LBNL on 31 GW of system performance.

14 SOLAR ENERGY↗

Data-Driven Mapping of the Cesium Cadmium Bromide Phase Space Utilizing a Soft-Chemistry Approach

Soft-chemistry techniques provide a versatile approach to synthesizing inorganic materials under mild conditions, enabling access to compositions and structures that are challenging to achieve through traditional thermodynamically driven solid-state methods. However, these solution-based routes often result in phase competition, requiring precise control over reaction conditions to achieve selective product formation. While one-variable-at-a-time (OVAT) approaches have traditionally been used for phase selection, data-driven strategies are emerging as more efficient methods for navigating complex synthetic spaces. Ternary metal halides, such as cesium cadmium bromides (Cs–Cd–Br), are of growing interest due to their potential in wide and ultrawide band gap applications. Unlike the well-studied cesium lead halide phases, the compositional diversity and solution-based synthesis of ternary Cs–Cd–Br phases remain largely unexplored. This study systematically investigates the synthetic phase space of the Cs–Cd–Br system by constructing a data-driven phase map. Using a common set of precursors and a standardized experimental procedure, we successfully synthesize all four known Cs–Cd–Br phases—CsCdBr 3 , Cs 2 CdBr 4 , Cs 3 CdBr 5 , and Cs 7 Cd 3 Br 13 —each exhibiting distinct structures, morphologies, and optical properties. Our findings highlight the potential of soft-chemistry methods for expanding the library of ternary metal halides and provide key insights into the thermodynamic and kinetic factors governing phase formation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

HyRAM+ (Hydrogen Plus Other Alternative Fuels Risk Assessment Models) v.4.0

HyRAM+ is a software toolkit for conducting quantitative risk assessment (QRA) and consequence modeling for hydrogen and other alternative fuels infrastructure and transportation systems. HyRAM+ contains validated, simplified release behavior models, a standardized QRA approach, and engineering models and generic data relevant to hydrogen installations. HyRAM (hydrogen-only) versions 1.0 to 3.1 were developed by Sandia for the U.S. Department of Energy (DOE) Hydrogen and Fuel Cell Technologies Office (HFTO). The U.S. DOE Vehicle Technologies Office (VTO) and U.S. Department of Transportation (DOT) Pipeline and Hazardous Material Safety Administration (PHMSA) contributed to the development of HyRAM+ version 4.0 regarding the addition of methane (natural gas) and propane models. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525. SAND2021-10862 O

Groth, Katrina↗

HyRAM+ (Hydrogen Plus Other Alternative Fuels Risk Assessment Models) v.4.1.1

HyRAM+ is a software toolkit for conducting quantitative risk assessment (QRA) and consequence modeling for hydrogen and other alternative fuels infrastructure and transportation systems. HyRAM+ contains validated, simplified release behavior models, a standardized QRA approach, and engineering models and generic data relevant to hydrogen installations. HyRAM (hydrogen-only) versions 1.0 to 3.1 were developed by Sandia for the U.S. Department of Energy (DOE) Hydrogen and Fuel Cell Technologies Office (HFTO). The U.S. DOE Vehicle Technologies Office (VTO) and U.S. Department of Transportation (DOT) Pipeline and Hazardous Material Safety Administration (PHMSA) contributed to the development of HyRAM+ version 4.0 regarding the addition of methane (natural gas) and propane models. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525. SAND2021-10862 O

Groth, Katrina↗

HyRAM+ (Hydrogen Plus Other Alternative Fuels Risk Assessment Models) v.5.1.1

HyRAM+ is a software toolkit for conducting quantitative risk assessment (QRA) and consequence modeling for hydrogen and other alternative fuels infrastructure and transportation systems. HyRAM+ contains validated, simplified release behavior models, a standardized QRA approach, and engineering models and generic data relevant to hydrogen installations. HyRAM (hydrogen-only) versions 1.0 to 3.1 were developed by Sandia for the U.S. Department of Energy (DOE) Hydrogen and Fuel Cell Technologies Office (HFTO). The U.S. DOE Vehicle Technologies Office (VTO) and U.S. Department of Transportation (DOT) Pipeline and Hazardous Material Safety Administration (PHMSA) contributed to the development of HyRAM+ version 4.0 regarding the addition of methane (natural gas) and propane models. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525. SAND2021-10862 O

Groth, Katrina↗

Low-Carbon District Heating: Performance Modeling of Hybrid Solar, Heat Pump, and Thermal Storage Systems for District Thermal Energy in the United States

District heating requires thermal energy in the temperature range of 40 degrees C - 120 degrees C. Typically, the thermal energy input for these systems has largely been met through fossil energy. However, the temperature range is low enough that it presents an opportunity for low-carbon technologies such as solar thermal and electrified thermal generators like heat pumps to decarbonize the heat generation. In this paper, a heat pump model was applied to estimate the performance and economics of a real-world low-carbon district heating substation. This system is comprised of a flat plate solar collector field paired with a mechanical vapor compression heat pump and hot water thermal storage, augmented by gas-fired boilers. Plant data was used to tune the model and estimate the system's benefits in terms of both standard financial metrics (IRR and payback), and environmental metrics, including avoided CO2 emissions. The model is subsequently employed to estimate the technical and economic potential of solar + heat pump + ther-mal storage hybrid systems as retrofits for district heating systems in eight US Markets.

district heating↗

Comparative Assessment of Data-driven Process Models in Health Information Technology

Process mining for conformance analysis consists of comparing a reference process model against a data-driven process model generated via log files from information technology systems. However, in the absence of a complete reference process model, we found no suggested approaches in the literature to address the need for evaluating process conformance among different healthcare facilities to assess standardization of care. Our goal is to find similarities and dissimilarities in data-driven process models among US Veterans Health Administration (VHA) facilities that can be indicative of patient safety issues. Our hypothesis was that the analysis would not produce statistically significant differences in outcome. We present a unique implementation of conformance analysis in process mining that consists of combining process mining, process mapping and statistical metrics. We illustrate our approach by applying it to the analysis of two clinical radiology order process models generated from healthcare data provided by two similar facilities in the VHA. The comparative assessment showed that about 70% of the orders completed successfully and 30% were not completed due to policy and duplications. Our analysis found a good statistical correlation between both facilities, as the Spearman’s correlation coefficient between facilities for the frequency of cases per total hours was 0.87879, for the frequency of cases by state transition was 0.79702 and for the throughput time per state transition was 0.63582. Additional statistical analyses using the Mann-Whitney U test and the root mean square error both produced values that were not significant. The foregoing approach validated our hypothesis by demonstrating a good statistical correlation of data describing the flow of clinical radiology orders absent a credible reference model. Finding good agreement between both facilities was important in confirming that the clinical orders flow in a similar manner, suggesting standardization of care.

97 MATHEMATICS AND COMPUTING↗

Empirical Comparison of Machine Learning Approaches for Black-Box Modeling of Power Conversion System Dynamics

Inverter-based resources are key components in modern power systems, but accurately modeling their complex behavior can be challenging. Standard, generic converter models often oversimplify inverter dynamics, leading to significant errors in predicting performance. In this work, we compare several data-driven machine learning (ML) approaches for inverter modeling, performing experiments on power conversion systems, systematically varying input conditions, and recording the resulting voltages and currents. The ML models were then trained on this measured data to capture the inverter's dynamic response and to predict the inverter's output current. A performance comparison between the four ML models under study is conducted, laying the foundation for future work on hardware implementation for real-time inference.

30 DIRECT ENERGY CONVERSION↗

Mixed-precision iterative refinement using tensor cores on GPUs to accelerate solution of linear systems

Double-precision floating-point arithmetic (FP64) has been the de facto standard for engineering and scientific simulations for several decades. Problem complexity and the sheer volume of data coming from various instruments and sensors motivate researchers to mix and match various approaches to optimize compute resources, including different levels of floating-point precision. In recent years, machine learning has motivated hardware support for half-precision floating-point arithmetic. A primary challenge in high-performance computing is to leverage reduced-precision and mixed-precision hardware. We show how the FP16/FP32 Tensor Cores on NVIDIA GPUs can be exploited to accelerate the solution of linear systems of equations Ax = b without sacrificing numerical stability. The techniques we employ include multiprecision LU factorization, the preconditioned generalized minimal residual algorithm (GMRES), and scaling and auto-adaptive rounding to avoid overflow. We also show how to efficiently handle systems with multiple right-hand sides. On the NVIDIA Quadro GV100 (Volta) GPU, we achieve a 4×-5× performance increase and 5× better energy efficiency versus the standard FP64 implementation while maintaining an FP64 level of numerical stability.

GMRES↗

Empirical Validation of UBEM: An Assessment of Bias in Urban Building Energy Modeling for Chicago

Residential and commercial buildings currently account for 30% of total global final energy consumption. Urban-scale building energy modeling (UBEM) can enable scalable investments and unlock building improvements by quantifying energy, demand, emissions, and cost reductions of specific measures or packages for building-specific technologies in large geographic regions. While the sophistication of UBEM data sources and technologies have increased dramatically in the past decade, there remains a knowledge gap for empirical validation and sources of bias between building-specific energy models and measured data at varying geographic scales.As UBEM continues to develop, systemic analysis of accuracy, bias, and limitations of the resulting models is necessary to inform best practices and move toward standardization. These are characterized for the Automatic Building Energy Modeling (AutoBEM) software suite with an initial case study involving metered electricity consumption data from 247,188 buildings in Chicago, Illinois, USA - averaged across years 2019-2021 - compared to the following datasets: (1) the AutoBEM-generated nation-scale Model America version 2 (MAv2) data for 596,064 buildings, (2) tax assessor data for 579,829 buildings, (3) tax assessor data filled with MAv2, and (4) 102 representative dynamic archetypes. The accuracy is reported for every building type and vintage combination, along with multiple sources of bias for unique building descriptors. The AutoBEM simulation workflow produced energy consumption estimates that closely match aggregated metered electricity consumption data for different types of buildings constructed during various time periods at the city scale - with initial normalized mean bias error of 10.9%, and 1.1% after removing outliers. Contribution of statistically significant factors including building type, land use, age, and size to variance in UBEM bias is quantified.

Garg, Ankur↗