Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “research data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Distilling Knowledge from Ensembles of Cluster-Constrained-Attention Multiple-Instance Learners for Whole Slide Image Classification

The peculiar nature of whole slide imaging (WSI), digitizing conventional glass slides to obtain multiple high resolution images which capture microscopic details of a patient’s histopathological features, has garnered increased interest from the computer vision research community over the last two decades. Given the unique computational space and time complexity inherent to gigapixel-size whole slide image data, researchers have proposed novel machine learning algorithms to aid in the performance of diagnostic tasks in clinical pathology. One effective algorithm represents a Whole slide image as a bag of smaller image patches, which can be represented as low-dimension image patch embeddings. Weakly supervised deep-learning methods, such as cluster-constrained-attention multiple instance learning (CLAM), have shown promising results when combined with image patch embeddings. While traditional ensemble classifiers yield improved task performance, such methods come with a steep cost in model complexity. Through knowledge distillation, it is possible to retain some performance improvements from an ensemble, while minimizing costs to model complexity. In this work, we implement a weakly supervised ensemble using clustering-constrained-attention multiple-instance learners (CLAM), which uses attention and instance-level clustering to identify task salient regions and feature extraction in whole slides. By applying logit-based and attention-based knowledge distillation, we show it is possible to retain some performance improvements resulting from the ensemble at zero cost to model complexity.

Alamudun, Folami↗

ARM Aerial Facility (AAF) Merged Value-Added Product Report for Historical G-1 Field Campaigns

For 30 years, the U.S. Department of Energy (DOE) Office of Science supported an instrumented Grumman Gulfstream-1 (G-1) aircraft for atmospheric field campaigns. Data from the final decade of G-1 operations were archived by the Atmospheric Radiation Measurement (ARM) user facility Data Center (ADC) and made publicly available at no cost to all registered users. To ensure a consistent data format and to improve the accessibility of the ARM airborne data, an integrated data set was recently developed covering the final six years of G-1 operations (2013 to 2018). The integrated data set includes data collected from 236 flights (766.4 hours). Four of the seven field campaigns were based in the U.S. One campaign collected data from the wildfires in the U.S. Pacific Northwest and agricultural burns in the lower Mississippi River valley as part of the Biomass Burning Observation Project (BBOP) in 2013. In 2015, the ARM Cloud Aerosol Precipitation Experiment provided data on atmospheric rivers and associated aerosol-cloud interactions that produce heavy precipitation on the U.S. west coast during the early spring. Research data from Airborne Carbon Measurements-V (ACME-V), collected during the summer of 2015, gave scientists insight into trends and variability of trace gases in the atmosphere over the North Slope of Alaska to improve arctic climate models. In the early summer and autumn of 2016, the Holistic Interactions of Shallow Clouds, Aerosols, and Land-Ecosystems (HI-SCALE) campaign provided an extensive data set geared toward coupled processes that affect the life cycle of shallow clouds through the interaction among aerosol, cloud, land surface, and ecosystems. In 2014 (March and October), the airborne sampling moved outside of the U.S. to the city of Manaus in central Amazonia, Brazil, where residential and industrial emissions were extensively characterized by flights of the G-1. The GoAmazon2014/15 aircraft campaign data are being integrated with aquatic and terrestrial ecosystem measurements to quantify anthropogenic perturbations to a usually pristine tropical environment. Another international airborne mission was carried out in the Eastern North Atlantic region. The Aerosol and Cloud Experiments in the Eastern North Atlantic (ACE-ENA) campaign saw the G-1 aircraft fly from Terceira Island in the Azores during the summer of 2017 and the winter of 2018. The campaign studied both seasons to measure key aerosol and cloud processes under various meteorological and cloud conditions with different aerosol sources. Then the G-1 deployed to the Sierras de Córdoba range in central Argentina from October to November 2018 for the Cloud, Aerosol, and Complex Terrain Interactions (CACTI) campaign to study orographic convective cloud interactions with their surrounding environment. These comprehensive datastreams provide much-needed insight into spatiotemporal variability of thermodynamic quantities, aerosol and cloud states, and properties for addressing essential science questions in Earth system process studies.

54 ENVIRONMENTAL SCIENCES↗

ARM Aerial Facility (AAF) Merged Value-Added Product Report for Historical G-1 Field Campaigns

For 30 years, the U.S. Department of Energy (DOE) Office of Science supported an instrumented Grumman Gulfstream-1 (G-1) aircraft for atmospheric field campaigns. Data from the final decade of G-1 operations were archived by the Atmospheric Radiation Measurement (ARM) user facility Data Center (ADC) and made publicly available at no cost to all registered users. To ensure a consistent data format and to improve the accessibility of the ARM airborne data, an integrated data set was recently developed covering the final six years of G-1 operations (2013 to 2018). The integrated data set includes data collected from 236 flights (766.4 hours). Four of the seven field campaigns were based in the U.S. One campaign collected data from the wildfires in the U.S. Pacific Northwest and agricultural burns in the lower Mississippi River valley as part of the Biomass Burning Observation Project (BBOP) in 2013. In 2015, the ARM Cloud Aerosol Precipitation Experiment provided data on atmospheric rivers and associated aerosol-cloud interactions that produce heavy precipitation on the U.S. west coast during the early spring. Research data from Airborne Carbon Measurements-V (ACME-V), collected during the summer of 2015, gave scientists insight into trends and variability of trace gases in the atmosphere over the North Slope of Alaska to improve arctic climate models. In the early summer and autumn of 2016, the Holistic Interactions of Shallow Clouds, Aerosols, and Land-Ecosystems (HI-SCALE) campaign provided an extensive data set geared toward coupled processes that affect the life cycle of shallow clouds through the interaction among aerosol, cloud, land surface, and ecosystems. In 2014 (March and October), the airborne sampling moved outside of the U.S. to the city of Manaus in central Amazonia, Brazil, where residential and industrial emissions were extensively characterized by flights of the G-1. The GoAmazon2014/15 aircraft campaign data are being integrated with aquatic and terrestrial ecosystem measurements to quantify anthropogenic perturbations to a usually pristine tropical environment. Another international airborne mission was carried out in the Eastern North Atlantic region. The Aerosol and Cloud Experiments in the Eastern North Atlantic (ACE-ENA) campaign saw the G-1 aircraft fly from Terceira Island in the Azores during the summer of 2017 and the winter of 2018. The campaign studied both seasons to measure key aerosol and cloud processes under various meteorological and cloud conditions with different aerosol sources. Then the G-1 deployed to the Sierras de Córdoba range in central Argentina from October to November 2018 for the Cloud, Aerosol, and Complex Terrain Interactions (CACTI) campaign to study orographic convective cloud interactions with their surrounding environment. These comprehensive datastreams provide much-needed insight into spatiotemporal variability of thermodynamic quantities, aerosol and cloud states, and properties for addressing essential science questions in Earth system process studies.

54 ENVIRONMENTAL SCIENCES↗

First Report of the Nuclear Data Subcommittee of the Nuclear Science Advisory Committee

Accurate, reliable nuclear data is essential for the success of Federal missions such as nonproliferation, nuclear forensics, homeland security, national defense, space exploration, clean energy generation, and scientific research. Data access is also key to innovative commercial developments such as new medicines, automated industrial controls, energy exploration, energy security, nuclear reactor design, and isotope production. The United States Nuclear Data Program (USNDP) is the domestic custodian of nuclear data. In its April 2022 meeting, the DOE/NSF Nuclear Science Advisory Committee was charged with preparing two reports on nuclear data. In this first report, we review recent accomplishments of the USNDP and discuss complementary and collaborative international efforts. Detailed descriptions of nuclear data needs for basic science, nonproliferation, national security, nuclear energy together with medical and space applications are also presented. Lastly, a set of specific cross-cutting nuclear data needs with relevance for multiple applications areas are also identified for further discussion in a follow-on report planned for release at the end of January 2023.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Computational tools and data integration to accelerate vaccine development: challenges, opportunities, and future directions

The development of effective vaccines is crucial for combating current and emerging pathogens. Despite significant advances in the field of vaccine development there remain numerous challenges including the lack of standardized data reporting and curation practices, making it difficult to determine correlates of protection from experimental and clinical studies. Significant gaps in data and knowledge integration can hinder vaccine development which relies on a comprehensive understanding of the interplay between pathogens and the host immune system. In this review, we explore the current landscape of vaccine development, highlighting the computational challenges, limitations, and opportunities associated with integrating diverse data types for leveraging artificial intelligence (AI) and machine learning (ML) techniques in vaccine design. We discuss the role of natural language processing, semantic integration, and causal inference in extracting valuable insights from published literature and unstructured data sources, as well as the computational modeling of immune responses. Furthermore, we highlight specific challenges associated with uncertainty quantification in vaccine development and emphasize the importance of establishing standardized data formats and ontologies to facilitate the integration and analysis of heterogeneous data. Through data harmonization and integration, the development of safe and effective vaccines can be accelerated to improve public health outcomes. Looking to the future, we highlight the need for collaborative efforts among researchers, data scientists, and public health experts to realize the full potential of AI-assisted vaccine design and streamline the vaccine development process.

60 APPLIED LIFE SCIENCES↗

FAIRmaterials: Ontology Tools with Data FAIRification in Development

The bilingual FAIRmaterials package simplifies the creation and visualization of materials and data science ontologies. FAIRmaterials, available in the Python and R languages, addresses the complexities associated with traditional ontology editors based on manual user input such as Protege with an intuitive workflow and easy-to-use templates, making it accessible to users both experienced and inexperienced with ontologies. The FAIRmaterials package is its ability to programatically convert simple and structured CSV inputs into rich, well-defined ontologies. This capability is designed to support the findability, accessibility, interoperability, and reusability (FAIR) of research data and serve as a tool in the process of data FAIRification. Its additional features, such as automated ontology merging, static visualizations, and comprehensive documentation for outputs extend its utility, making it a valuable tool for any researcher engaged in knowledge management.

Bradley, Alexander Harding [Case Western Reserve U↗

2025 Workshop on Envisioning Frontiers in AI and Computing for Biological Research: Position Papers

This workshop aims to identify key research directions for transforming biology using artificial intelligence (AI), machine learning (ML) and computational methods to facilitate the discovery of new behaviors, mechanisms, and designs of biological processes relevant to DOE missions, underpinning a broader U.S. bioeconomy. By developing novel AI/ML technologies to analyze and interpret complex biological data, researchers can organize and simulate biological processes at various scales as well as advance predictive understanding and manipulation of biological systems. This integration of computation, experimentation, and next-generation experimental technologies can lead to discoveries in new biological behaviors and mechanisms relevant to DOE missions. The focus is on how advanced computational and mathematical methods can impact this mission by exploring digital twins, foundation models, automated laboratory experiments, modeling of complex living systems, and data-driven approaches for the biodesign of plants and microbial systems. While data management is important, it is not the primary focus of this workshop, which will assess the current state, trends, and AI/ML challenges at the interface between biology and computational science to identify opportunities for high-impact research at their intersection. The goal is to define research needs and opportunities that align with biological sciences, computational sciences, and applied mathematics research.

59 BASIC BIOLOGICAL SCIENCES↗

Data and Scripts associated with a manuscript on ecosystem responses to wildfires in the Columbia River Basin

This data package is associated with the publication “Ecosystem leaf area, gross primary production, and evapotranspiration responses to wildfire in the Columbia River Basin” submitted to Biogeosciences (Shi et al., 2024; doi: 10.22541/au.171053013.30286044/v1). In this research, data products, leaf area index (LAI), gross primary production (GPP), and evapotranspiration (ET), from the Moderate Resolution Imaging Spectroradiometer (MODIS) are used to quantify the resistance and resilience of different ecosystem types in the Columbia River Basin (CRB). A machine learning algorithm, random forest (RF), was used to examine the impacts of precipitation, vapor pressure deficit (VPD), and burn severity from Monitoring Trends in Burn Severity (MTBS) on ecosystem resilience. The data package includes the processed MODIS data products, precipitation, VPD, and burn severity in 138 fire regions in CRB and the input files for RF model training. This data package includes six folders. The MODIS products are included in three MODIS_* folders with shell scripts for data clipping and *ncl files for data processing: (1) “/MODIS_LAI_CRB”; (2) “/MODIS_GPP_CRB”; and (3) “/MODIS_ET_CRB”. All the processed data for each fire event are NetCDF formatted. The MTBS burn severity data and the shell and *ncl scripts used for data processing are in the folder named (4) “MTBS_fire”. The ERA meteorological fields and the data processing scritps are in (5) “ERA_Var_CR”. All the scripts for figure development are in the format of *ncl and in the folder (6) “paper_scripts”. See the file ending in “flmd.csv” for a list of all files contained in this data package and descriptions for each. Tabular column headers and units are described in the data dictionary file ending in “dd.csv”.

54 ENVIRONMENTAL SCIENCES↗

Web-based Preprocessing and Visualization of 3D FIB Tomography Data for Nuclear Fuel Characterization

Three-dimensional (3D) focused ion beam (FIB) tomography enables reconstruction of internal nuclear fuel features that can't be fully evaluated through surface imaging alone. This capability supports characterization of fuel constituents and defects under thermal and irradiation conditions relevant to microreactor development. However, large tomography datasets can create data-handling, loading, and visualization challenges, especially when image-stack preparation and file conversion must be completed with separate tools. The Computational Ultraspatial Tomography Toolkit for High-Resolution Object Analysis Tools (CUTTRHOAT) is an open-source web application being developed to display FIB tomography datasets available through the Nuclear Research Data System (NRDS). The current alpha version requires prepared HDF5 datasets and has limited integrated data-preparation capabilities. This project improves CUTTHROAT by adding dataset-folder selection, automatic input detection, dataset scanning, missing-slice identification, blank-slice insertion, and image-stack-to-HDF5 conversion. Two applications will be compared: the baseline CUTTHROAT alpha workflow and the updated application containing the integrated data-handling and preprocessing functions. Evaluation will consider dataset detection accuracy, conversion success, loading time, rendering responsiveness, application stability, and user interaction. Preliminary results demonstrate successful loading of existing HDF5 files and converted image stacks, while testing also identified performance reductions caused by excessive blank-slice generation. The updated workflow reduces reliance on external preparation tools and supports more direct movement from image stacks to color-code 3D visualization. Future work includes refining missing-slice handling, integrating additional preprocessing functions, like a denoising feature, parsing TIFF metadata for automatic voxel scaling, and adding manual X, Y, and Z voxel-spacing inputs for PNG and JPEG.

36 - MATERIALS SCIENCE↗

V1G Frequency Regulation: Algorithm Development, Validation & Analysis at Scale

Researchers at Argonne National Laboratory developed and validated a high-fidelity digital twin of a smart charging (V1G) ecosystem to model the participation of up to 1,000 unique electric vehicles (EVs) in the PJM frequency regulation market. Utilizing a discrete-event framework, the simulation models complex interactions, from dynamic grid signals (updated every 2 seconds) to individual EV charging dynamics. The simulation incorporates multiple EV models created from real-world lab test data. Researchers tested multiple control algorithms to balance the dual objectives of maximizing aggregator’s revenue and driver charging needs. Results demonstrate that aggregated EVs function as a controllable, highly effective grid resource, achieving high PJM Performance Scores (80–90%). Additionally, an optimized, market-aware bidding strategy was identified as key to profitability. The platform was shown to provide drivers with an average charging discount of nearly 50%. The algorithm was further validated in the lab using production EVs and charging stations to compare simulation results with real-world performance.

Manne, Nithin↗

An open source knowledge graph ecosystem for the life sciences

Translational research requires data at multiple scales of biological organization. Advancements in sequencing and multi-omics technologies have increased the availability of these data, but researchers face significant integration challenges. Knowledge graphs (KGs) are used to model complex phenomena, and methods exist to construct them automatically. However, tackling complex biomedical integration problems requires flexibility in the way knowledge is modeled. Moreover, existing KG construction methods provide robust tooling at the cost of fixed or limited choices among knowledge representation models. PheKnowLator (Phenotype Knowledge Translator) is a semantic ecosystem for automating the FAIR (Findable, Accessible, Interoperable, and Reusable) construction of ontologically grounded KGs with fully customizable knowledge representation. The ecosystem includes KG construction resources (e.g., data preparation APIs), analysis tools (e.g., SPARQL endpoint resources and abstraction algorithms), and benchmarks (e.g., prebuilt KGs). We evaluated the ecosystem by systematically comparing it to existing open-source KG construction methods and by analyzing its computational performance when used to construct 12 different large-scale KGs. With flexible knowledge representation, PheKnowLator enables fully customizable KGs without compromising performance or usability.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Recent advances in plasma control and physics research in the Large Helical Device

The Large Helical Device (LHD), the largest superconducting helical system in the world, is equipped with advanced heating and diagnostic tools, facilitating plasma control and physics research. Data assimilation was employed for electron temperature control using a real-time Thomson scattering system and real time prediction code. A virtual LHD environment enabled visualization of escaping high-energy tritium ions and demonstrated that these ions impact the rear side of the divertor plate. Pioneering results crucial to plasma control have also been achieved. Real-time wall conditioning using Lithium granule dropping improved bulk ion energy and particle transport while simultaneously enhancing the heavy impurity transport. Progress has also been made in the investigation of turbulence-driven transport. At the confinement bifurcation, ion-scale turbulence decreased, while electron-scale turbulence increased. A change in the anisotropy of turbulent eddies was also observed at the confinement bifurcation. Coexistence of local and non-local turbulence was identified in electron-scale turbulence. Non-local turbulence exhibited the rapid spatial propagation of perturbations throughout the plasma, while local turbulence followed the temperature gradient. A transition between drift-wave turbulence and magnetohydrodynamics (MHD) turbulence was observed with the turbulence minimized at the transition condition. Machine learning analysis was employed to evaluate the temperate and density conditions of this turbulence transition. Then, real-time control of fueling and heating was applied to maintain the turbulence transition condition, improving the energy confinement enhancement factor by 20%. In addition, evidence was obtained for collisionless ion heating by energetic-ion-driven geodesic acoustic modes and MHD bursts. These achievements represent unique contributions to the development of fusion reactors.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A labeled dataset for building HVAC systems operating in faulted and fault-free states

Abstract Open data is fueling innovation across many fields. In the domain of building science, datasets that can be used to inform the development of operational applications - for example new control algorithms and performance analysis methods - are extremely difficult to come by. This article summarizes the development and content of the largest known public dataset of building system operations in faulted and fault free states. It covers the most common HVAC systems and configurations in commercial buildings, across a range of climates, fault types, and fault severities. The time series points that are contained in the dataset include measurements that are commonly encountered in existing buildings as well as some that are less typical. Simulation tools, experimental test facilities, and in-situ field operation were used to generate the data. To inform more data-hungry algorithms, most of the simulated data cover a year of operation for each fault-severity combination. The data set is a significant expansion of that first published by the lead authors in 2020.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Bridging New Observational Capabilities and Process-Level Simulation: Insights into Aerosol Roles in the Earth System

The spatial distribution of ambient aerosol particles significantly impacts aerosol–radiation–cloud interactions, which contribute to the largest uncertainty in global anthropogenic radiative forcing estimations. However, the atmospheric boundary layer and lower free troposphere have not been adequately sampled in terms of spatiotemporal resolution, hindering a comprehensive characterization of various atmospheric processes and impeding our understanding of the Earth system. To address this research data gap, we have leveraged the development of uncrewed aerial systems (UAS) and advanced measurement techniques to obtain mesoscale spatial data on aerosol microphysical and optical properties around the U.S. Southern Great Plains (SGP) atmospheric observatory. Our study also benefits from state-of-the-art laboratory facilities that include three-dimensional molecular imaging techniques enabled by secondary ion mass spectrometry and nanogram-level chemical composition analysis via micronebulization aerosol mass spectrometry. Through our study, we have developed a framework for observation–modeling integration, enabling an examination of how various assumptions about the organic–inorganic components mixing state, inferred from chemical analysis, affect clouds and radiation in observation-constrained model simulations. By integrating observational constraints (derived from offline chemical analysis of the aerosol surface using collected samples) with in situ UAS observations, we have identified a prominent role of organic-enriched nanometer layers located at the surface of aerosol particles in determining profiles of aerosol optical and hygroscopic properties over the SGP observatory. Furthermore, we have improved the agreement between predicted clouds and ground-based cloud lidar measurements. This UAS–model–laboratory integration exemplifies how these new advanced capabilities can significantly enhance our understanding of aerosol–radiation–cloud interactions.

54 ENVIRONMENTAL SCIENCES↗

Standardizing Scientific Metadata at Los Alamos National Laboratory

This document is intended for publication in Descriptive Notes, the blog of the Description Section of the Society of American Archivists. This paper focuses on cataloging processes for scientific datasets and the creation of scientific metadata to ensure accessibility of LANL researchers' data.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Wireless High-Temperature Sensor Network for smart boiler systems

This final project report describes the research data and findings. This project aims to develop a new wireless high-temperature sensor network for real-time continuous boiler condition monitoring in harsh environments. Such a wireless high-temperature sensor network enables network-based automatic temperature sensing and data collection, which combined with artificial intelligent (AI) algorithms allow the construction of smart boiler systems with boiling condition management and optimization for significant energy-saving and reliability improvement

42 ENGINEERING↗

It's All About the Envelope: Prioritizing Envelope Upgrades for Electrification of Cold Climate Homes

Building decarbonization via electrification on a clean grid is the most promising climate solution proposed to date for the building sector. In cold climate zones, building electrification will be driven in large part by moving from natural gas space heating to cold climate heat pumps (CCHPs). CCHPs are commercially available today, including economical cold climate air source heat pumps (ccASHPs). But there's one big problem - wide-scale adoption of ccASHPs will dramatically increase winter peak electricity demand, even with the highest efficiency ccASHP products. Furthermore, cold climate space heating loads will drive unprecedented electric system peaks during the lowest periods of renewable generation and are likely to overwhelm existing distribution systems. This scenario is avoidable by coupling electrification with building envelope upgrades to reduce peak heating loads. This paper presents a model, built from home energy audit and research data sets, that quantifies the above challenges. Results demonstrate how weatherization efforts coupled with additional high-performance envelope upgrade measures can prepare the building stock for electrification and show the benefit these measures can bring to future utility operations. Much of this envelope upgrade work is cost-effective, according to conservative cost-benefit testing and program successes to date, and is coupled with substantial non-energy benefits. However, persistent market barriers have made scaling of envelope retrofit work challenging for decades, suggesting additional policy support is required. Lessons learned from previous policy experience, combined with new technology and administrative support, create exciting potential for this decarbonization climate solution.

air sealing↗