Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “research data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

What can simulation test beds teach us about social science? Results of the ground truth program

The ground truth program used simulations as test beds for social science research methods. The simulations had known ground truth and were capable of producing large amounts of data. This allowed research teams to run experiments and ask questions of these simulations similar to social scientists studying real-world systems, and enabled robust evaluation of their causal inference, prediction, and prescription capabilities. We tested three hypotheses about research effectiveness using data from the ground truth program, specifically looking at the influence of complexity, causal understanding, and data collection on performance. We found some evidence that system complexity and causal understanding influenced research performance, but no evidence that data availability contributed. The ground truth program may be the first robust coupling of simulation test beds with an experimental framework capable of teasing out factors that determine the success of social science research.

97 MATHEMATICS AND COMPUTING↗

Freewheeling: What Six Locations, 61,000 Trips, and 242,000 Miles in Colorado Reveal about How E-Bikes Improve Mobility Options

Personal micromobility modes such as bicycles, e-bikes and scooters offer low- or zero-emission transportation alternatives to single occupancy vehicles (SOVs). However, the lack of supporting data has led to a dearth of data-driven research on the usage of personally owned e-bikes, including variations due to weather, geography and demographics. In this paper, we present an overview of the longitudinal findings from the CanBikeCO program, focused on e-bike adoption and use rates across different demographics, trip characteristics, and geographies. The CanBikeCO program recorded travel survey data from late July 2021 to December 31, 2022, from low-income Colorado households who were provided with e-bikes for personal use by the Colorado Energy Office (CEO). This data was collected in six different communities across Colorado following the mini-pilot program that was conducted in Fall 2020. To collect data for the survey, the program used the NREL OpenPATH application, which combines passive data collection with semantic information such as trip mode and purpose labels. To the best of our knowledge, there is no prior travel survey data on personally owned e-bikes with this range and scope. This unique dataset yielded several insights. One is that commute trips among participants had nearly 17% higher shares of e-bikes than all trips combined. E-bikes were stated to most often replace cars (34% of e-bike trips) and personal micromobility (22%). Participants favored walking for trips less than 1 mile, e-bikes for trips 1-3 miles, and e-bikes, cars or shared rides for trips 3-20 miles. Seasonality accounted for a 10% decrease and subsequent recovery in e-bike mileage on a per user basis. E-bikes are also appealing across age groups, even among older individuals, and see decreased utilization similar to regular bikes or walking during winter months. We also find that e-bike use may be related to characteristics of land use and urban form, occupation and income as well as household car ownership. We conclude that, for this population, who are mainly part of low-income households, the emissions added by the use of e-bikes (in the case of replacement of non-motorized modes) are outweighed by the strong single occupancy vehicle (SOV) travel replacement. As a whole, our findings suggest a considerable potential for energy savings and emissions reductions from personal e-bike ownership.

ADVANCED PROPULSION SYSTEMS↗

PSInet: a new global water potential network

Abstract Given the pressing challenges posed by climate change, it is crucial to develop a deeper understanding of the impacts of escalating drought and heat stress on terrestrial ecosystems and the vital services they offer. Soil and plant water potential play a pivotal role in governing the dynamics of water within ecosystems and exert direct control over plant function and mortality risk during periods of ecological stress. However, existing observations of water potential suffer from significant limitations, including their sporadic and discontinuous nature, inconsistent representation of relevant spatio-temporal scales and numerous methodological challenges. These limitations hinder the comprehensive and synthetic research needed to enhance our conceptual understanding and predictive models of plant function and survival under limited moisture availability. In this article, we present PSInet (PSI—for the Greek letter Ψ used to denote water potential), a novel collaborative network of researchers and data, designed to bridge the current critical information gap in water potential data. The primary objectives of PSInet are as follows. (i) Establishing the first openly accessible global database for time series of plant and soil water potential measurements, while providing important linkages with other relevant observation networks. (ii) Fostering an inclusive and diverse collaborative environment for all scientists studying water potential in various stages of their careers. (iii) Standardizing methodologies, processing and interpretation of water potential data through the engagement of a global community of scientists, facilitated by the dissemination of standardized protocols, best practices and early career training opportunities. (iv) Facilitating the use of the PSInet database for synthesizing knowledge and addressing prominent gaps in our understanding of plants’ physiological responses to various environmental stressors. The PSInet initiative is integral to meeting the fundamental research challenge of discerning which plant species will thrive and which will be vulnerable in a world undergoing rapid warming and increasing aridification.

Forestry↗

Green Computing Opportunities & Strategy

Computation is critical to emerging fields of data intensive research, enabling new methodologies, approaches and tools. As the rate of hardware efficiency gains slows, computational time and energy costs increase. To match pace with computation demand, new approaches are needed to keep the opportunity for impact open. Charles Tripp, lead of the Green Computing Catalyzer, discusses research efforts to improve software efficiency to enable faster, less energy-intensive computing.

algorithmic efficiency↗

Advances in scientific literature mining for interpreting materials characterization

Abstract Using synchrotron light sources, such as the National Synchrotron Light Source II at Brookhaven National Laboratory, scientists in fields as diverse as physics, biology, and materials science, identify the atomic structure, chemical composition, or other important properties of varied specimens. x-ray spectroscopy from light sources is particularly valuable for materials research with vast information available about reference spectra in the scientific literature. However, as the technique is applicable to many science domains, searching for information about select x-ray spectroscopy spectra is impeded by the sheer number of publications. Moreover, useful information about the context of an experiment or figures presented in papers can be buried among the details, which takes time to assess. This work presents a scientific literature mining system that supports data acquisition, information extraction, and user interaction for referencing x-ray spectra identification and spectral interpretation. The goal is to provide efficient access to useful spectral data to researchers who may spend only a few days at a synchrotron light source. With this system, users browse a classification tree for papers arranged according to x-ray spectroscopic methods, chemical elements, and x-ray absorption spectroscopy edges. Relevant figures are extracted with sentences from the paper that explain them, known as ‘figure explanatory text.’ Notably, this system focuses on semantic aspects (logical analysis) to find figure explanatory text using deep contextualized word embeddings techniques and contains an interface to obtain labeled data from domain experts that is used to evaluate and improve the model.

Park, Gilchan (ORCID:0000000201536646)↗

Audi e-tron Green Light Optimized Speed Advisory On-Road Data

To aid researchers in studying the capabilities and benefits of vehicle-to-infrastructure communication, Argonne National Laboratory collected a robust set of on-road driving data of the Audi Green Light Optimized Speed Advisory (GLOSA) system implemented in the e-tron battery electric vehicle. This dataset includes 33 tests, each roughly 27 miles in length and roughly 45 to 75 minutes in duration. The team selected Kane County Highway Route 34 from Main Street in Batavia, Illinois to Middlecreek Lane in St. Charles, Illinois as the route do to its high density of GLOSA-active lights and the most opportunities to observe the system per hour of test time. The data include parameters from the following sources: GLOSA system driving the dash indicators, multiple powertrain parameters including real-time battery power/energy consumption, GPS, front radar gap, and rear radar gap. ![audio-e-tron image](audi-e-tron.jpg)

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Benchmarking universal machine learning interatomic potentials for rapid analysis of inelastic neutron scattering data

The accurate calculation of phonons and vibrational spectra remains a significant challenge, requiring highly precise evaluations of interatomic forces. Traditional methods based on the quantum description of the electronic structure, while widely used, are computationally expensive and demand substantial expertise. Emerging universal machine learning interatomic potentials (uMLIPs) offer a transformative alternative by employing pre-trained neural network surrogates to predict interatomic forces directly from atomic coordinates. This approach dramatically reduces computation time and minimizes the need for technical knowledge. In this paper, we produce a phonon database comprising nearly 5000 inorganic crystals to benchmark the performance of several leading uMLIPs. We further assess these models in real-world applications by using them to analyze experimental inelastic neutron scattering data collected on a variety of materials. Through detailed comparisons, we identify the strengths and limitations of these uMLIPs, providing insights into their accuracy and suitability for fast calculations of phonons and related properties, as well as the potential for real-time interpretation of neutron scattering spectra. Our findings highlight how the rapid advancement of AI in science is revolutionizing experimental research and data analysis.

inelastic neutron scattering↗

Genesis Mission Data cards

As data-intensive research and artificial intelligence become central to DOE mission science, the need for machine-actionable dataset documentation has grown accordingly. However, many DOE-aligned communities, including the Office of Science, NNSA, and cross-laboratory collaborations, have developed independent metadata practices. This fragmentation creates friction for discovery, federation, and reuse across programs. To address these challenges, this talk introduces the Genesis Data Card: a shared metadata artifact developed in collaboration with a broad DOE community (Jefferson Lab and the National Lab of the Rockies, Oak Ridge, Sandia, Idaho, Berkeley, and Los Alamos). The Genesis Data Card aims to standardize dataset documentation across DOE-aligned initiatives while remaining extensible to discipline-specific needs. This talk will describe the data card template and the supporting code to validate completed data cards, using a companion LinkML schema. I'll walk through the design decisions behind the template, its alignment with existing standards, its treatment of sensitivity and governance metadata, and the phased roadmap toward lifecycle-integrated "xCards" that support autonomous discovery and reuse. The talk closes with current gaps, ongoing work, and how others can contribute datasets and feedback to the shared repository.

McSpadden, Helen [Thomas Jefferson National Accele↗

Temperature, Humidity, and Time-Lapse Video Data from the East River Watershed, Water Year 2024

A new version of this dataset is available at doi:10.15485/3001338 and is the first citation in the 'Related References' section. It is expands on this dataset by appending another water year of data collection and additional logger sites.This dataset contains time-lapse imagery and distributed measurements of air temperature, relative humidity, dew point, and soil temperature across the East River basin from 3 October 2023 to 12 August 2024. Instruments were deployed at 14 sites as part of the DOE Grant: Seasonal Cycles Unravel Mysteries of Missing Mountain Water organized by Jessica Lundquist (University of Washington), Rosemary Carroll (Desert Research Institute), and Ethan Gutmann (National Center for Atmospheric Research). The data are published to support studies of surface climate or hydrologic processes in complex terrain. Measurements were collected with low-cost data loggers installed 2 m high on evergreen trees or buried just below the soil surface. Time-lapse cameras were deployed at three sites. Imagery from sites AP BONUS and AP5 provides insight into large-scale seasonal snow cover variability. Imagery from site EL2 shows smaller-scale snow patterns across a nearby meadow.Dataset files are organized by site and variable (air measurements, ground measurements, or time-lapse video). Air and ground measurements are packaged in LoggerData.zip, and time-lapse imagery is compiled into short videos stored in TimelapseVideos.zip. File-level metadata contains details for each file included in the dataset. A data dictionary provides units and descriptions for column or row names in all files. The locations metadata file describes site characteristics, locations, and associated GPS methods.Dataset update 2025-03-03: Resolved header and datetime formatting inconsistencies within LoggerData.zip files KP1_Air, KP3_Air, KP3_Ground, KP6_Ground, AP3_Ground, AP4_Air, and AP6_Air.Dataset update 2025-11-18: Modified abstract and related references sections to include new version of dataset.

54 ENVIRONMENTAL SCIENCES↗

GROWdb US River Systems - Samples

GROW Overview We developed the Genome Resolved Open Watersheds database (GROWdb), which aims to increase genomic sampling and understanding of global river microbiomes. An emphasis of GROWdb is to create a publicly available and ever-expanding microbial genome database that is focused on rivers while being interoperable with databases from other ecosystems. GROWdb is based on a network-of-networks approach to move beyond a small collection of well-studied rivers, towards a spatially distributed, global network of systematic observations. GROWdb represents the first microbial, river-focused resource parsed at various scales from genes to MAGs to community level including expression and potential based measurements that will be of interest to microbiologists, ecologists, geochemists, hydrologists, and modelers. Dataset Acknowledgement GROWdb contains data from various research campaigns, please acknowledge the following data generators, as appropriate: WHONDRS derived genomes or samples - include this statement in your acknowledgements: “This study used data from the Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS) under the River Corridor Science Focus Area (SFA) at the Pacific Northwest National Laboratory (PNNL) that was generated at the U.S. Department of Energy (DOE) Joint Genome Institute User Facility. PNNL is operated by Battelle Memorial Institute for the U.S. DOE under Contract No. DE-AC05-76RL01830. The SFA is supported by the U.S. DOE, Office of Biological and Environmental Research (BER), Environmental System Science (ESS) Program.” Total Samples loaded onto this Narrative: 178 Note: Not all GROW samples may be loaded into KBase Data Availability The data underlying GROWdb are accessible across various platforms to ensure all levels of data structure are widely available. First, all reads and MAGs are publicly hosted on National Center for Biotechnology (NCBI) under Bioproject PRJNA946291. Second, all data related data presented here including MAG annotations, extended data tables, phylogenetic tree files, antibiotic resistance gene database files, and MAG abundance tables are available in Zenodo (link). Beyond the flat database files listed above, our aim for GROWdb was to maximize data use by making the data available in searchable and interactive platforms including the National Microbiome Data Collaborative (NMDC) data portal, the Department of Energy’s Systems Biology Knowledgebase (KBase), and a GROW specific user interface released here, GROWdb Explorer. Each platform provides different ways to interact with GROWdb: NMDC GROWdb formed a pilot project for the NMDC. Specifically, individual GROWdb datasets (metagenomes, metatranscriptomes, etc) are easily accessible and searchable through the NMDC data portal, where they are systematically connected to each other and to a rich suite of sample information and standard analysis results, following Findable, Accessible, Interoperable, and Reusable (FAIR) data practices. KBase GROWdb is publicly available within KBase, including samples (this Narrative), MAGs, and corresponding genome scale metabolic models. Access within KBase allows for immediate access and reuse of data, including comparison to private data using KBase’s 500+ analysis tools. Other linked narratives in KBase: GROW Metagenome Assembled Genomes (MAGs) GROW Metabolic Models GROWdb Explorer GROWdb data is also explorable through a graphical user interface built through the Colorado State University Geospatial Centroid (https://geocentroid.shinyapps.io/GROWdatabase/), allowing users to search and graph microbial and spatial data simultaneously. In summary, this microbial genome resource represents the first publicly available genome collection from rivers and offers data that can be leveraged across microbiome studies. GROWdb is an expanding repository to incorporate and unify global river multi-omic data for the future.

59 BASIC BIOLOGICAL SCIENCES↗

Materials Data Science Ontology(MDS-Onto): Unifying Domain Knowledge in Materials and Applied Data Science

Ontologies have gained popularity in the scientific community as a way to standardize terminologies in organizations’ data. Although certain cohorts have created frameworks with rules and guidelines on creating ontologies, there exist significant variations in how Materials Science ontologies are currently developed. We seek to provide guidance in the form of a unified automated framework for developing interoperable and modular ontologies for Materials Data Science that simplifies the ontology terms matching by establishing a semantic bridge up to the Basic Formal Ontology(BFO). This framework provides key recommendations on how ontologies should be positioned within the semantic web, what knowledge representation language is recommended, and where ontologies should be published online to boost their findability and interoperability. Two fundamental components of the MDS-Onto framework are the bilingual package called FAIRmaterials for ontology creation and FAIRLinked, for FAIR data creation. To showcase the practical capabilities of FAIRmaterials, we present two exemplar domain ontologies of MDS-Onto: Synchrotron X-Ray Diffraction and Photovoltaics.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Involving the new generations in Fermilab endeavors

Since 1984 the Italian groups of the Istituto Nazionale di Fisica Nucleare (INFN) and Italian Universities, collaborating with the DOE laboratory of Fermilab (US) have been running a two-month summer training program for Italian university students. While in the first year the program involved only four physics students of the University of Pisa, in the following years it was extended to engineering students. This extension was very successful and the engineering students have been since then extremely well accepted by the Fermilab Technical, Accelerator, and Scientific Computing Division groups. Over the many years of its existence, this program has proven to be the most effective way to engage new students in Fermilab endeavors. Many students have extended their collaboration with Fermilab with their Master’s Thesis and PhD. Since 2004 the program has been supported in part by DOE in the frame of an exchange agreement with INFN. Over its almost 40 years of history, the program has grown in scope and size and has involved more than 550 Italian students from more than 20 Italian Universities, Several Institutes of Research, including ASI and INAF in Italy, and the ISSNAF Foundation in the US, have provided additional financial support. Since the program does not exclude appropriately selected non-Italian students, a handful of students from European and non-European Universities were also accepted over the years. Each intern is supervised by a Fermilab Mentor responsible for performing the training program. Training programs spanned from Tevatron, CMS, Muon (g-2), Mu2e, and Short Baseline Neutrino Experiments and DUNE design and experimental data analysis, development of particle detectors (silicon trackers, calorimeters, drift chambers, neutrino and dark matter detectors), design of electronic and accelerator components, development of infrastructures and software for exascale data handling, research on superconductive elements and on accelerating cavities, and theory of particle accelerators. Since 2010, within an extended program supported by the Italian Space Agency and the Italian National Institute of Astrophysics, a total of 30 students in physics, astrophysics, and engineering have been hosted for two months in the summer at US space science Research Institutes and laboratories. In 2015 the University of Pisa included these programs within its educational programs. Accordingly, Summer School students are enrolled at the University of Pisa for the duration of the internship and are identified and ensured as such. At the end of the internship, the students are required to write summary reports on their achievements. After positive evaluation by a University Examining Board, interns are acknowledged credits for their Diploma Supplement. The program was canceled in 2020 and 2021 due to the pandemic but restarted successfully in 2022. We believe this program can be taken as a model and easily adopted by interested institutions.

99 GENERAL AND MISCELLANEOUS↗

Survey of Time Shift Detection Algorithms for Measured PV Data

In this research, three variations of time shift detection algorithms were tested for their ability to detect time shift issues (including daylight savings time and random time shifts) in measured PV data sets. Two algorithms from the Python PVAnalytics package were assessed, and one algorithm from the Solar-Data-Tools package was assessed. Each algorithm's ability to accurately detect and measure time shifts was assessed.

automated preprocessing↗

Automating Bug Report Classification with Few Shot Learning

Orthogonal defect classification (ODC) is a method used to categorize software defects, providing valuable insights into the development process. This study focuses on automating the classification of software bug reports into different ODC defect types using few shot learning, a machine learning approach that requires minimal labeled data. Previous research has manually classified bug reports or used traditional machine learning algorithms like linear support vector machine, achieving limited success. Our approach uses few shot learning to improve classification accuracy and efficiency. The results show a harmonic mean of recall and precision (i.e., the F1 score) of around 0.6 which is a performance improvement over previous methods. The results highlight the potential benefit of few shot learning techniques and their application in enhancing the safety and reliability of nuclear digital instrumentation and control (DI&C) systems. Future work will explore incorporating advanced techniques to supplement the model's training data and achieve better results.

42 - ENGINEERING↗

An Overview of Argonne’s Advanced Mobility Technology Laboratory Vehicle Systems Instrumentation and Evaluation Methodology

This report will provide a general overview of the testing facilities, research equipment, and general testing methodologies Argonne utilizes to conduct vehicle technology evaluations of advanced technology research vehicles. Data is captured from vehicle evaluations that provide powertrain operation and corresponding energy consumption based on a combination of in-depth instrumentation and focused testing. This resulting dataset of hundreds of time-resolved vehicle signals provide a basis for direct analysis and model validation of vehicle specific technologies. Argonne has attempted to standardize the approach to vehicle instrumentation and testing at Argonne, but it should be noted that each vehicle, and the corresponding assessment, remains unique. It is suggested that the reader reference vehicle specific testing reports for an overview of the unique aspects for each test vehicle. Additionally, the datasets for research vehicles are made publicly available through the Advanced Mobility Technology Laboratory’s Downloadable Driving Database (D3) at www.anl.gov/d3.

33 ADVANCED PROPULSION SYSTEMS↗

Washington Air-Sea Interaction Research Facility Waves and Currents Test Data

This archive includes data from the University of Washington WASIRF (Washington Air-Sea Interaction Research Facility) flume. WASIRF is a laboratory testing tank at the Northwest National Marine Renewable Energy Center designed to investigate wind-wave-current interactions. It includes test data for simultaneous waves and current generation done at the WASIRF lab. A report included in the archive further details testing methodology. Wave and current data is provided in .dat files.

16 TIDAL AND WAVE POWER↗

FAIR principles for AI models with a practical application for accelerated high energy diffraction microscopy

Abstract A concise and measurable set of FAIR (Findable, Accessible, Interoperable and Reusable) principles for scientific data is transforming the state-of-practice for data management and stewardship, supporting and enabling discovery and innovation. Learning from this initiative, and acknowledging the impact of artificial intelligence (AI) in the practice of science and engineering, we introduce a set of practical, concise, and measurable FAIR principles for AI models. We showcase how to create and share FAIR data and AI models within a unified computational framework combining the following elements: the Advanced Photon Source at Argonne National Laboratory, the Materials Data Facility, the Data and Learning Hub for Science, and funcX, and the Argonne Leadership Computing Facility (ALCF), in particular the ThetaGPU supercomputer and the SambaNova DataScale ® system at the ALCF AI Testbed. We describe how this domain-agnostic computational framework may be harnessed to enable autonomous AI-driven discovery.

47 OTHER INSTRUMENTATION↗

Identifying Heterogeneous Micromechanical Properties of Biological Tissues via Physics–Informed Neural Networks

The heterogeneous micromechanical properties of biological tissues have profound implications across diverse medical and engineering domains. However, identifying full-field heterogeneous elastic properties of soft materials using traditional engineering approaches is fundamentally challenging due to difficulties in estimating local stress fields. Recently, there has been a growing interest in data-driven models for learning full-field mechanical responses, such as displacement and strain, from experimental or synthetic data. However, research studies on inferring full-field elastic properties of materials, a more challenging problem, are scarce, particularly for large deformation, hyperelastic materials. Here, a physics-informed machine learning approach is proposed to identify the elasticity map in nonlinear, large deformation hyperelastic materials. This study reports the prediction accuracies and computational efficiency of physics-informed neural networks (PINNs) in inferring the heterogeneous elasticity maps across materials with structural complexity that closely resemble real tissue microstructure, such as brain, tricuspid valve, and breast cancer tissues. Further, the improved architecture is applied to three hyperelastic constitutive models: Neo-Hookean, Mooney Rivlin, and Gent. Furthermore, the improved network architecture consistently produces accurate estimations of heterogeneous elasticity maps, even when there is up to 10% noise present in the training data.

59 BASIC BIOLOGICAL SCIENCES↗