Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “standardized data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Characterizing peak electricity demand for U.S. households: an assessment of end-use loads and demand factors

Understanding household peak electricity demand is critical to evaluate the technical need for electrical infrastructure upgrades. This study characterizes peak loads for existing and new equipment using metered data from a convenience sample of 11,940 U.S. dwellings from four sources, including 911 from two sources with end-use metering. After standardized data cleaning and labeling, we derived descriptive statistics for key metrics, such as maximum demand and demand factors, and developed predictive models relating 60- to 15-min demand for the National Electrical Code (NEC). Mean 15-min maximum demand was 9.7 kW (median 9.0 kW; IQR 7.0–11.5 kW, 95% CI 9.6–9.8 kW), indicating spare capacity in 98% of homes with hypothetical 100 A panels. Maximum demand increased with floor area and number of high-demand loads. Dwelling maximum demand was driven by higher-power, longer-duration heating appliances and vehicle charging, while most user-operated appliances contributed little. Demand factors are used to account for how most devices contribute less than their rated power to maximum demand. Existing load mean demand factors (28%; median 10%; IQR 0–58%; CI 28–29%) were higher than those for new loads (21%; median 7%; IQR 0–35%; CI 20–21%), because new loads changed the timing and magnitude of maximum demand. New high-demand loads had higher than average demand factors (40–60%). Whole dwelling demand factors support the NEC's 40% assumption, but they challenge its conservative 100% treatment of new HVAC. We propose a data-driven 50% demand factor for new equipment, which would align with metered data, improve affordability, and modernize electrical codes.

Appliances↗

Reactor Containment Passive Safety Analysis: Steam Condensation in Presence of Non-condensable Gas Scaled Experiment and Modeling

This study presents steam condensation scaled experiments and semi-empirical models in presence of nitrogen (N)—a noncondensable gas (NCG), simulating air in the reactor containment—to support water-cooled small modular reactors (SMRs) passive containment cooling system (PCCS) design and analysis. Previous experimental studies on PCCS are focused on fixed and smaller tube (mostly 2-in.) geometries and specific test condition variations, bringing challenges with geometric scaling and mismatching with SMR prototypic design. To address these challenges, this study presents steam condensation test dataset obtained from three scaled test sections of 1-, 2-, and 4-in.-diameter steam condensers with an annular/jacket cooling of 2-, 3-, and 6 in.-diameter tubes, respectively. Test data were collected for steam ranges from 58 to 63 kg/hr., and NCG flow of 4.4 to 13.3 kg/hr. Annular cooling water flow was varied to obtain required testing conditions of saturated steam inlet and fully condensed outlet. Axial temperature test data of bulk cooling water, steam and condensate were collected by thermocouples for three test sections and various steam-NCG mixing/testing conditions. A standard data reduction method was adopted—utilizing iterative and nodalized mass and heat transfer calculation—to estimate axial local heat fluxes, heat transfer coefficients (HTCs), condensation rates, film thickness, and Nusselt number. Based on the obtained dataset semi-empirical model results—a ratio of experimental and Nusselt’s theoretical HTC are presented. Such results and findings are supportive of developing scaled-up testing facility, to enable model validations and accelerate next generation of reactors development and deployment

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Presentation: Reactor Containment Passive Safety Analysis: Steam Condensation in Presence of Non-condensable Gas Scaled Experiment and Modeling

This study presents steam condensation scaled experiments and semi-empirical models in presence of nitrogen--a noncondensable gas (NCG), simulating air in the reactor containment--to support water-cooled small modular reactors (SMRs) passive containment cooling system (PCCS) design and analysis. Previous experimental studies on PCCS are focused on fixed and smaller tube (mostly 2-in.) geometries and specific test condition variations, bringing challenges with geometric scaling and mismatching with SMR prototypic design. To address these challenges, this study presents steam condensation test dataset obtained from three scaled test sections of 1-, 2-, and 4-in.-diameter steam condensers with an annular/jacket cooling of 2-, 3-, and 6 in.-diameter tubes, respectively. Test data were collected for steam ranges from 58 to 63 kg/hr., and NCG flow of 4.4 to 13.3 kg/hr. Annular cooling water flow was varied to obtain required testing conditions of saturated steam inlet and fully condensed outlet. Axial temperature test data of bulk cooling water, steam and condensate were collected by thermocouples for three test sections and various steam-NCG mixing/testing conditions. A standard data reduction method was adopted--utilizing iterative and nodalized mass and heat transfer calculation to estimate axial local heat fluxes, heat transfer coefficients (HTCs), condensation rates, film thickness, and Nusselt number. Based on the obtained dataset semi-empirical model results--a ratio of experimental and Nusselt's theoretical HTC are presented. Such results and findings are supportive of developing scaled-up testing facility, to enable model validations and accelerate next generation of reactors development and deployment.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Livewire Data Platform: File Standards Version 1.0

This technical document is a user guide to help users of the Livewire Data Platform understand the standards and requirements for storing and sharing data on the Livewire Data Platform.

97 MATHEMATICS AND COMPUTING↗

POWTEX visits POWGEN

The high-intensity time-of-flight (TOF) neutron diffractometer POWTEX for powder and texture analysis is currently being built prior to operation in the eastern guide hall of the research reactor FRM II at Garching close to Munich, Germany. Because of the world-wide 3 He crisis in 2009, the authors promptly initiated the development of 3 He-free detector alternatives that are tailor-made for the requirements of large-area diffractometers. Herein is reported the 2017 enterprise to operate one mounting unit of the final POWTEX detector on the neutron powder diffractometer POWGEN at the Spallation Neutron Source located at Oak Ridge National Laboratory, USA. As a result, presented here are the first angular- and wavelength-dependent data from the POWTEX detector, unfortunately damaged by a 50 g shock but still operating, as well as the efforts made both to characterize the transport damage and to successfully recalibrate the voxel positions in order to yield nonetheless reliable measurements. Also described is the current data reduction process using the PowderReduceP2D algorithm implemented in Mantid [Arnold et al. (2014). Nucl. Instrum. Methods Phys. Res. A , 764 , 156–166]. The final part of the data treatment chain, namely a novel multi-dimensional refinement using a modified version of the GSAS-II software suite [Toby & Von Dreele (2013). J. Appl. Cryst. 46 , 544–549], is compared with a standard data treatment of the same event data conventionally reduced as TOF diffraction patterns and refined with the unmodified version of GSAS-II . This involves both determining the instrumental resolution parameters using POWGEN's powdered diamond standard sample and the refinement of a friendly-user sample, BaZn(NCN) 2 . Although each structural parameter on its own looks similar upon comparing the conventional (1D) and multi-dimensional (2D) treatments, also in terms of precision, a closer view shows small but possibly significant differences. For example, the somewhat suspicious proximity of the a and b lattice parameters of BaZn(NCN) 2 crystallizing in Pbca as resulting from the 1D refinement (0.008 Å) is five times less pronounced in the 2D refinement (0.038 Å). Similar features are found when comparing bond lengths and bond angles, e.g. the two N—C—N units are less differently bent in the 1D results (173 and 175°) than in the 2D results (167 and 173°). The results are of importance not only for POWTEX but also for other neutron TOF diffractometers with large-area detectors, like POWGEN at the SNS or the future DREAM beamline at the European Spallation Source.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A guide to the BRAIN Initiative Cell Census Network data ecosystem

Characterizing cellular diversity at different levels of biological organization and across data modalities is a prerequisite to understanding the function of cell types in the brain. Classification of neurons is also essential to manipulate cell types in controlled ways and to understand their variation and vulnerability in brain disorders. The BRAIN Initiative Cell Census Network (BICCN) is an integrated network of data-generating centers, data archives, and data standards developers, with the goal of systematic multimodal brain cell type profiling and characterization. Emphasis of the BICCN is on the whole mouse brain with demonstration of prototype feasibility for human and nonhuman primate (NHP) brains. Here, we provide a guide to the cellular and spatial approaches employed by the BICCN, and to accessing and using these data and extensive resources, including the BRAIN Cell Data Center (BCDC), which serves to manage and integrate data across the ecosystem. We illustrate the power of the BICCN data ecosystem through vignettes highlighting several BICCN analysis and visualization tools. Finally, we present emerging standards that have been developed or adopted toward Findable, Accessible, Interoperable, and Reusable (FAIR) neuroscience. The combined BICCN ecosystem provides a comprehensive resource for the exploration and analysis of cell types in the brain.

59 BASIC BIOLOGICAL SCIENCES↗

Surface Water Quality Data from Beaver-Impacted Streams; Trail Creek and East River, Colorado 2025

This data package contains surface water chemistry measurements collected in 2025 to evaluate how beaver damming and low-tech process-based stream restoration influence water quality and metal mobility in mountainous headwater systems of the Upper Colorado River Basin. Sampling was conducted at Trail Creek (Taylor Park watershed, Colorado), a tributary undergoing restoration through installation of low-tech process-based structures (i.e., beaver dam analogs), and at off-channel beaver ponds within the East River floodplain (East River watershed, Colorado). Samples were collected along longitudinal transects spanning upstream control reaches, beaver-influenced ponded reaches, and downstream segments. Additional samples were collected from near-surface pore waters within a beaver dam seepage face. The dataset includes concentrations of major and trace elements measured by inductively coupled plasma–mass spectrometry (ICP-MS) and inductively coupled plasma–optical emission spectrometry (ICP-OES), major anions measured by ion chromatography (IC), and dissolved organic carbon (DOC; reported as non-purgeable organic carbon, NPOC). Samples were size-fractionated at 0.45 micrometers (µm), 0.22 µm, and 0.02 µm to distinguish particulate (>0.45 µm), colloidal (0.22–0.02 µm), and dissolved (<0.02 µm) fractions. The data package consists of comma-separated value (.csv) files containing tabulated chemical concentration data, sample metadata (site identifiers, geographic coordinates, sampling dates, fraction type), and quality control flags. All files are provided in open, non-proprietary formats that can be accessed using standard data analysis software such as Microsoft Excel, R, Python, MATLAB, or other programs capable of reading .csv files. Units, detection limits, and analytical methods are documented in accompanying metadata files. The dataset is designed to support analyses of (1) how beaver impoundment and restoration structures alter elemental partitioning and transport, (2) the role of iron and organic carbon in mediating trace metal mobility, and (3) reach-scale changes in water quality across restoration gradients. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

Anions↗

Methods for safely sharing dual-use genetic data

Background: Some genetic data has dual-use potential. Sharing pathogen data has shown tremendous value. For example therapeutic development and lineage tracking during the COVID pandemic. This data sharing is complicated by the fact that these data have the potential to be used for harm. The genome sequence of a pathogen can be used to enable malicious genetic engineering approaches or to recreate the pathogen from synthetic DNA. Standard data security methods can be applied to genetic data, but when data is shared between institutions, ensuring appropriate security can be difficult. Sensitive data that is shared internationally among a wide array of institutions can be especially difficult to control. Methods for securely storing and sharing genetic data with potential for dual-use are needed to mitigate this potential harm.Results: Here we propose new methods that allow genetic data to be shared in a data format that prevents a nefarious actor from accessing sensitive aspects of the data. Our methods obfuscate raw sequence data by pooling reads from different samples. This approach can ensure that data is secure while stored and during electronic transfer. We demonstrate that by pooling raw sequence data from multiple samples of the same organism, the ability to fully reconstruct any individual sample is prevented. In the pooled data, most genomic information remains, but reads or mutations cannot be directly attributed to any individual sample. To further restrict access to information, regions of a genome can be removed from the reads.Conclusion: Our methods obscure genomic information within raw sequence reads. This method can allow genetic data to be stored and shared while preventing a nefarious actor from being able to perfectly reconstruct an organism. Broad-scale sequence information remains, while fine scale details about specific samples are difficult or impossible to reconstruct. Our software is available at https://github.com/Geneinfosec-Inc/ReadMixer.

59 BASIC BIOLOGICAL SCIENCES↗

Improving Residential Building Simulations Through Large-Scale Empirical Validation

Residential building energy simulations are increasingly used for energy-efficient building design, codes and standards analysis, home certifications and ratings, utility programs, and technology assessments. Various software tools exist to perform residential building simulations, and these tools often use different models, inputs, and assumptions. This leads to inconsistencies that can undermine confidence in the predicted results. Validation of these tools can increase confidence by ensuring their accuracy and consistency. One way to validate simulation tools is through empirical testing, which compares predicted energy usage to measured utility billing data. This paper describes the process of data collection, data standardization, and empirical validation, and illustrates its use with our residential EnergyPlus (R)-based software. The data and process can be extended to other simulation tools and contribute to improving residential building simulations more broadly.

empirical validation↗

NEPATEC2.0: NEPA Text Corpus v2.0

The National Environmental Policy Act of 1969, as amended (NEPA), is a major environmental law in the United States, requiring Federal agencies to consider and document potential environmental impacts before deciding on a proposed action. Modernization of NEPA and permitting processes faces significant challenges due to the lack of standardized formats and interoperable systems for organizing and sharing NEPA-related information across agencies. Much of the information gathered during NEPA reviews is written into documents such as categorical exclusions, environmental assessments, and environmental impact statements, then filed in predominately independent agency file stores that may or may not be publicly accessible. The application of metadata and data standards, such as those recommended by the Council on Environmental Quality (CEQ), to NEPA documents offers a shared vocabulary and structure for key entities like projects, processes, and documents that can streamline information exchange and enhance collaboration across systems. In this work, we publicly release NEPATEC2.0, an expanded corpus of NEPA documents with associated metadata. NEPATEC2.0 encompasses approximately 120,000 documents from 60,000 projects prepared by more than 60 different agencies. Modeled to align with CEQ metadata standards, NEPATEC2.0 promotes consistency in environmental reviews and supports the ongoing effort to modernize permitting technologies by facilitating more transparent, efficient, and data-driven decision-making. Importantly, NEPATEC2.0 demonstrates the possibilities and limitations of large language model-based prompting to extract information from NEPA documents at scale.

environmental review↗

NEPATEC v2.0: Standardized Metadata and Text Corpus of National Environmental Policy Act Documents

The National Environmental Policy Act of 1969, as amended (NEPA), is a major environmental law in the United States, requiring Federal agencies to consider and document potential environmental impacts before deciding on a proposed action. Modernization of NEPA and permitting processes faces significant challenges due to the lack of standardized formats and interoperable systems for organizing and sharing NEPA-related information across agencies. Much of the information gathered during NEPA reviews is written into documents such as categorical exclusions, environmental assessments, and environmental impact statements, then filed in predominately independent agency file stores that may or may not be publicly accessible. The application of metadata and data standards, such as those recommended by the Council on Environmental Quality (CEQ), to NEPA documents offers a shared vocabulary and structure for key entities like projects, processes, and documents that can streamline information exchange and enhance collaboration across systems. In this work, we publicly release NEPATEC2.0, an expanded corpus of NEPA documents with associated metadata. NEPATEC2.0 encompasses approximately 120,000 documents from 60,000 projects prepared by more than 60 different agencies. Modeled to align with CEQ metadata standards, NEPATEC2.0 promotes consistency in environmental reviews and supports the ongoing effort to modernize permitting technologies by facilitating more transparent, efficient, and data-driven decision-making. Importantly, NEPATEC2.0 demonstrates the possibilities and limitations of large language model-based prompting to extract information from NEPA documents at scale.

54 ENVIRONMENTAL SCIENCES↗

Integrated Energy-Water Data for Cross-Sector Resilience

This white paper focuses on the “energy-for-water” domain, addressing the urgent need for integrated, empirical data to support regional management, benchmarking, and research on improving efficiency and developing technologies for water and wastewater management systems. The costs and energy required for the supply, treatment, and distribution of water and wastewater lack a standard data collection mechanism and centralized database or storage infrastructure, limiting data-driven decision-making across interdependent infrastructure systems.

42 ENGINEERING↗

MPEX AI Digital Twins

All magnetically confined plasma fusion power plant concepts (Tokamak, Spherical Tokamak, Stellarator, Mirror, ...) must exhaust the heat and plasma from the core confinement region to the material walls. The primary channel for this exhaust is through a plasma divertor which directs plasma along open magnetic field lines to a material target. The Material Plasma Exposure eXperiment (MPEX) illustrated in Figure 1, is a high-power, steady-state linear plasma device designed to produce the plasma material interaction (PMI) conditions of the divertor of future magnetic confinement fusion power plants: energy flux 20MW/m 2 , ion fluence 1031/m 2 , pulse duration 106 sec. These goals of plasma exposure in MPEX are well beyond those achieved in magnetic fusion experimental devices. Successfully achieving these high power steady state conditions for long pulses requires operational control of the heating and particle sources and the plasma flux to the walls and target. The MPEX AI Hot Spot Controller, proposed in this project, will help achieve the operational milestones of MPEX. The MPEX device will begin commissioning at the end of FY26. A smaller proto-MPEX was operated for 14,666 plasma discharges and will resume operation in September of 2025 as proto-MPEX-lite, with reduced capability, to test a new window for the Helicon plasma source. The proto-MPEX data has undergone surrogate modeling with machine learning methods (R. Archibald, 2022 IEEE International Conference on Big Data). This proto-MPEX data will be used to begin development of the AI digital twins described in this white paper. The scientific mission of MPEX is to qualify materials of different composition for use in the high energy and plasma flux conditions of a fusion power plant. The materials exposed in MPEX will in some cases be exposed to high neutron fluxes at other ORNL facilities to measure the changes to their PMI properties. The targets exposed in MPEX will be transported under vacuum to a Surface Analysis Station (SAS). The SAS will be equipped with the following diagnostics: Focused Ion Beam (FIB) for trench milling, 100-400 angstrom resolution scanning electron microscope (SEM), surface mapping x-ray spectrometer, high resolution camera, and a future upgrade to a laser induced breakdown spectroscopy quadruple mass spectrometer (LIBS-QMS). The MPEX experiments will generate diverse pre- and post-exposure measurement data of detailed material properties down to the crystal grain level in 3D for post-exposure assessment of PMI damage (e.g. cracking, melting, erosion and redeposition of the material). Physics models for the PMI, and how the material composition and manufacturing impact its performance under high energy plasma exposure, need to be validated with MPEX data to guide the selection of new candidate materials. Our vision for the MPEX AI Digital Twins project is to supply experimental and physics model simulation data to train Artificial Intelligence (AI) models for data processing, analysis, operational control, PMI and materials simulation to maximize the scientific output of the MPEX device. Ultimately, an AI digital twin of MPEX material assessment metrics for tested and synthetic material types with simulated PMI will be trained by the AI Modeling Teams on the experimental and physics simulation data submitted to the American Science Cloud by this project. A purely empirical search for the best material is inefficient given the finite number of samples that can be tested on MPEX. In order to expand the material properties database for training the MPEX Material Assessment AI Digital Twin, and to gain physics understanding of the PMI processes, physics models of the material properties and PMI processes are required. The physics simulations provide detailed simulation data, like impact angles for plasma ions, sputtering yields, transport of the ionized sputtered target material in the plasma, and redeposition locations. This simulation data expands the measurement data for deeper physics understanding. The experimental data is essential to validate the PMI and material structure simulation models. The validated models can then be used to generate new simulation data of MPEX material assessments for synthetic material compositions that have not been exposed in MPEX. These predictive simulations, plus the whole experimental dataset, will be used to train the MPEX Material Assessment AI Digital Twin allowing a rapid generative AI search for new materials with reduced PMI damage by interpolating the domain of the training set. These new optimum materials can be simulated with the physics codes and/or tested in MPEX. The ability of AI neural networks to interpolate multi-dimensional parameter spaces and generate virtual data is exploited for a more efficient search for optimum materials. The advent of the Transformational AI Models Consortium (TAIMC) is an opportunity to engage with state of the art private and public AI developers to achieve the goals of the AI digital twins and AI accelerated physics models proposed in this project. Our partners at ORNL from the Advance Scientific Computing Research (ASCR) organization will collaborate in accelerating the integrated plasma material interaction simulation framework. This simulation framework will provide a platform for generating simulation data across a range of physical fidelities, including hybrid methods that produce multi-fidelity results. This data will be leveraged for AI model development, both for generation of surrogates and the automation of simulation campaigns. A part of the research below will include collaborative efforts with the TAIMC to (i) adapt data storage approaches to ensure AI-readiness, (ii) provide a protypical exemplar to inform and exercise constructed workflows, and (iii) generate and share data, using the TAIMC unified AI data standard, for foundational models that will be trained from multiple sources across the DOE complex. We will also collaborate with the TAIMC, as well as the planned AI modeling teams, to develop approaches for reducing the cost of data generation. These include tailored multi-fidelity approaches as well as fine-tuning strategies to augment general, large-scale foundational models.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Idaho National Laboratory Data Acquisition And Processing System

The INLDAS data acquisition and processing system is designed to develop prototype measurements and real-time processing techniques. The INLDAS hardware and firmware currently consists of National Instruments • NI-DAQmx 14.1 • cDAQ-9184 • 9205 • 9211 INLDAS can perform of standard data acquisition functions as well as novel functions and real-time processing algorithms. There are 3 acquisition modes to choose from: • Continuous mode runs when the user hits start, processing data and logging it to file • SWTrigger mode waits for one or more predefined triggers before acquiring data. It will buffer data as well, so it can record data that happened shortly before the trigger • Wakeup mode waits on predefined timers. When a time goes off, it acquires a preset amount of data There are also 3 Data Processing Modes • Normal mode does no processing besides the rolling average • FFT Mode Performs an FFT on incoming data every time the time window has elapsed • STFFT mode

Smith, JamesA↗

Comprehensive Evaluation of Agrivoltaics Research: Breadth, Depth, and Insights for Future Research

Agrivoltaics integrates agricultural production with solar energy generation to address challenges related to land use, food security, and renewable energy development. This study provides the most comprehensive evaluation to date of global agrivoltaic research, aiming to classify the literature, identify strengths and gaps, and guide future work. We systematically screened over 3000 English-language publications through 2023 for relevant agrivoltaic publications. A total of 670 studies were categorized in the InSPIRE Data Portal across five agrivoltaic activities and multiple hierarchical themes, including physical, biological, technological, social, and crosscutting domains. We found that research was concentrated on crop production, microclimate dynamics, and PV performance, with gaps in areas like human health, wildlife, policy, and standardized methodologies. Although the U.S. emphasizes animal grazing and habitat-based systems in practice, most U.S.-based studies focused disproportionately on crop production. The analysis revealed uneven geographic and topical representation and highlighted a lack of integrated, interdisciplinary approaches. This study concludes that while agrivoltaic research has grown rapidly, more coordinated efforts could support standardized data collection, address overlooked ecological and social impacts, and align research focus with real-world system implementation, ultimately improving the scalability and successful deployment of agrivoltaic systems.

14 SOLAR ENERGY↗

Dataset Repository for Investigating Suicide Risk Using Social and Environmental Determinants of Health

Suicide is frequently modeled as a function of genetics and environment, where the latter refers to factors other than direct biological consequences, such as air quality, financial level, social connectivity, transportation and food access, and homelessness status. According to the World Health Organization, clean air, a stable climate, adequate water, sanitation and hygiene, safe chemical use, radiation protection, healthy and safe workplaces, sound agricultural practices, health-supportive cities and built environments, and a preserved natural environment are all prerequisites for good health. Understanding the relationships between these determinants and mental health outcomes requires standardized data that can be included in healthcare programs and health outcome models. There is a wealth of publicly available data on social and environmental factors provided by various US organizations that can benefit the design of health care systems and public health interventions, as well as improve our comprehension of factors that impact health. Such information would not only help improve the understanding of individual and community risk but also identify new risk factors that have not previously been therapeutically targeted, especially in terms of their impact on mental health. However, curating and standardizing such datasets is challenging because they are often recorded at numerous geographical and temporal resolutions and with varying spatial and temporal granularities. To address this challenge, we launched an endeavor in conjunction with the Veterans Health Administration to collect publicly available socioeconomic and environmental determinants of health statistics in the US. In this manuscript, we describe a social and environmental determinants of health (SEDH) datasets repository, data curation documentation, and a pipeline framework for data generation; This effort started in 2020, when we began constructing a scalable pipeline to automate the download, extraction, preparation, analysis, and production of datasets. These datasets have been made available to the VHA and may be shared upon agreement with collaborating organizations.

60 APPLIED LIFE SCIENCES↗

Preserving nonlinear constraints in variational flow filtering data assimilation

Data assimilation aims to estimate the states of a dynamical system by optimally combining sparse and noisy observations of the physical system with uncertain forecasts produced by a computational model. The states of many dynamical systems of interest obey nonlinear physical constraints, and the corresponding dynamics is confined to a certain sub-manifold of the state space. Standard data assimilation techniques applied to such systems yield posterior states lying outside the manifold, violating the physical constraints. This work focuses on particle flow filters which use stochastic differential equations to evolve state samples from a prior distribution to samples from an observation-informed posterior distribution. The variational Fokker-Planck (VFP)—a generic particle flow filtering framework—is extended to incorporate non-linear, equality state constraints in the analysis. To this end, two algorithmic approaches that modify the VFP stochastic differential equation are discussed: (i) VFPSTAB, to inexactly preserve constraints with the addition of a stabilizing drift term, and (ii) VFPDAE, to exactly preserve constraints by treating the VFP dynamics as a stochastic differential-algebraic equation (SDAE). Additionally, an implicit-explicit time integrator is developed to evolve the VFPDAE dynamics. The strength of the proposed approach for constraint preservation in data assimilation is demonstrated on three test problems: the double pendulum, Korteweg-de-Vries, and the incompressible Navier-Stokes equations.

97 MATHEMATICS AND COMPUTING↗

A Data Science and Machine Learning Platform Supporting Large Particle Accelerator Control and Diagnostics Applications Final Report: SBIR Initial Phase II DE-SC0022583

The Machine Learning Data Platform (MLDP) is a product providing full-stack support for data science, Machine Learning, and Artificial Intelligence (ML/AI) applications at particle accelerator and large experimental physics facilities. It supports ML/AI applications from front-end, high-speed acquisition of heterogeneous, time-series data, through data archiving and management, to back-end analysis. The MLDP embodies a “data-science ready” platform for data analysis and ML/AI applications in diagnosis, modelling, control, and optimization of these facilities. It provides data scientists and applications a consistent, datacentric interface to archive data standardizing implementation and deployment of ML/AI algorithms to different operations configurations within the same facility, or between facilities. Being an open-source, public-domain project, the MLDP is intended for broadest possible impact by increasing accessibility and minimizing the required expertise for installation and operation. The MLDP can also be deployed at user facilities for experimental data collection, archiving, and analysis. It is capable of acquisition and archiving of heterogeneous data from experimental equipment (e.g., images, arrays, structures, etc.) along with system hardware configurations (e.g., scalars, tables), control system process variables, and any metadata required for provenance. Thus, the MLDP can manage experimental data through its entire lifecycle, from acquisition and archiving, through analysis and investigation, to release and final publication.

43 PARTICLE ACCELERATORS↗