Data Collection and Analysis [Slides]
The report summarizes the ongoing data collection as well as selection and analysis of field events based on measurement data from project partners.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
The report summarizes the ongoing data collection as well as selection and analysis of field events based on measurement data from project partners.
This presentation introduces INL's human reliability analysis data collection studies.
Energy-dispersive X-ray diffraction (EDXD) at synchrotron beamlines is commonly used for the study of material properties under high pressure and/or high temperature. Experimenters typically rely on the availability of robust data collection and analysis at a beamline, but this has become increasingly difficult, especially with the introduction of multi-element detectors that generate complex, multi-dimensional data sets. These data sets have energy resolution, and they can also be resolved in relation to sample position, diffraction angle, or different external stimuli. We report a new Python-based graphical program, hpMCA, for EDXD data collection and analysis that streamlines the experimental process for the beamline users. The program features a user-friendly interface, capability for online viewing and analyzing data from multi-element energy-dispersive detectors, and includes features useful for working with samples under high pressure and/or high temperature, such as crystal phase identification, real-time unit cell lattice refinement, and pressure determination based on an equation of state.
SF-25-043 This software package facilitates data collection for the purpose of assessing the performance of LLMs on scientific topics
This presentation presents a brief overview of the collection of information about the US government's fleet of motor vehicles using the Federal Automotive Statistical Tool (FAST), discusses the makeup and operation of the vehicle fleet during FY 2020, discusses challenges associated with quality of the submitted data, and touches on future aspects of fleet data collection and reporting. FAST is a web-based information system sponsored by GSA's Office of Government-wide Policy and DOE's Federal Energy Management Program to collect information about the US federal government's fleet of motor vehicles; FAST is developed, maintained, and supported by DOE's Idaho National Laboratory (INL).
The Midwest Regional Carbon Initiative (MRCI) Task 3.0 was defined to facilitate development of carbon capture, utilization, and storage (CCUS) in the region by collection and sharing of existing and new technical data from CCUS projects and research. The task also included support for further analysis and assessment of tools by the project team and by researchers working on programs such as National Risk Assessment Partnership (NRAP), machine learning (ML) techniques, and assessment and improvement of CCUS site assessment, operations, and monitoring aspects. Work under Task 3.0 addressed key issues related to CCUS deployment and provided foundational research and datasets to help establish CCUS projects in the MRCI. Report Authors and Principal Technical Contributors: Joel Sminchak, Laura Keister, Mackenzie Scharenberg, Priya Ravi-Ganesh, Autumn Haagsma, Srikanta Mishra, Jared Hawkins, Jared Schuetter, Amy Lang, Jaelen Lewis, Derrick James, Jorge Barrios, Stuart Skopec, and Sanjay Mawalkar (Battelle). Chris Korose, Carl Carmen, Nate Grigsby, Nathan Webb (Illinois State Geological Survey). Principal Investigators: Dr Neeraj Gupta, Dr. Chris Korose.
This research introduces an energy prediction framework at the facility level supported by automated data collection and machine learning models. It investigates whether reducing the prediction time scale allows for applying more complex machine learning techniques and if those techniques improve the prediction accuracy. The primary advantages of this framework lie in its automation of the energy prediction process and its provision of real-time energy data suitable for use in energy dashboards or digital twins. A sitewide dataset was created by combining 15 min energy and daily production data of five shops—assembly, battery, body (electric), body (gas), and paint—from a globally recognized electric vehicle manufacturer. Various machine learning models were evaluated on daily, weekly, and monthly datasets, including, in increasingly complex order: naïve, simple linear regression, net regularized generalized linear regression, principal component regression, k-nearest neighbor, random forest, and Bayesian regularized neural network. Compared to the current state-of-the-art energy consumption prediction for the industrial facility level, this research investigates more complex models and smaller time intervals for higher accuracy. The findings revealed that the more complex monthly models require a minimum of a year and a half of data to operate, while weekly models demand a year of data to achieve improved accuracy. Daily models can operate with only six months of data but exhibit poor performance due to reduced prediction accuracy of production. Key challenges identified include access to reliable, high-quality energy and production data and the initial demand for human labor.
During the COVID-19 pandemic of 2020, major case reporting outlets quickly coalesced around two or three primary vendors. Johns Hopkins University and The New York Times were among the more prominent, and all were of great value to the nation, particularly during the uncertain early stages of the pandemic. They primarily focused on three major attributes: number of new cases, deaths, and recovery, but only at the state level. Recognizing that many states were reporting very detailed data sets (e.g., hospital beds) at a count level or finer, the ORNL Pandemic Modeling team embarked on a major data curation effort from March to June 2020 for the purpose of capturing this wealth of detailed data. The challenge of curating this data was daunting. The number of attributes reported by the states grew on almost on a weekly basis. States were routinely shifting their web tool strategies away from easily parsable HTML-based formatting to new Tableau and ArcGIS content. This growth in the sheer number of attributes combined with the unpredictable shifts in data format meant an aggressive and agile combination of automated scripting and manual scraping was required to capture new daily streams. To keep up, the team had to scale up staff and widen its approach for capture and storage. The DOE COVID-19 data collection effort resulted in over 11 million data points being collected, covering over 13,000 unique geographies and over 2,000 unique attributes that spanned predominantly from early March through the end of June 2020.
The "data collection" basically involves setting up the thermal wave frequency, laser scan distance, and other parameters related to the experimental setup. The modification of this code is minor and the details of this code can be found in the earlier patent ("thermal conductivity microscope"). The "data analysis" instead, replaces the simplified analytical model by a more complete analytical model, and used a "thermoquadruple" method to solve the analytical model. The efficiency is orders of magnitude improved and the accuracy is also better. Meanwhile, the previous model can only handle a two-layer sample structure. The new, complete model can handle materials with multiple layers (any given number), which is necessary to handle post ion irradiated materials.
Come learn about the new Hydrogen Component Reliability Database (HyCReD) and participate in discussions on hydrogen component reliability data collection, collaboration, and analysis. Funded by the U.S. Department of Energy's Office of Energy Efficiency and Renewable Energy under the Hydrogen and Fuel Cell Technologies Office, HyCReD is a collaborative project between the National Renewable Energy Laboratory, the University of Maryland, and hydrogen stakeholders to improve safety reliability for hydrogen facilities by integrating risk reduction methodologies and component reliability data taxonomies that support hydrogen infrastructure failure rate analysis.
In this project, we assembled a sensor suite that combines a scientific echosounder (sonar) system with video and acoustic cameras (secondary sensors). The sensor suite generates data that is amenable to automated target detection algorithms and can provide inputs to animal encounter models. We developed software (archiving software) to collect data simultaneously from all sensors in the suite, analyze the sonar data automatically in near real-time to identify time periods when targets of interest were present, and automatically archive data from the secondary sensors for these time periods. The product of the archiving software is a data set for the secondary sensors, containing data only for the times when targets of interest were determined to be present by the sonar. We performed controlled field testing to verify the operation of the sensor suite and the archiving software.
The US Department of Energy (DOE) Oak Ridge Reservation (ORR) is located in Anderson and Roane Counties, Tennessee. A portion of the ORR, known as Self-Sufficiency Parcel 2 (SSP2) is planned for transfer for private use. The SSP2 Site is approximately 670 acres (Figure 1-1), although the current plan is to only clear and develop a portion of this acreage. Any inquiries about the land transfer and future development should be directed to DOE Oak Ridge Environmental Management, as this is beyond the scope of the Natural Resources Management Team (NRMT). NRMT records bat data for the entire ORR, including acoustic monitoring, mist netting and cave surveys. A few surveys have previously been conducted for small land transfers adjacent to SSP2 (See Appendix A), but not for the entire SSP2 area. Since bats have a large range, it was decided that collecting data while SSP2 was still accessible would be beneficial for the NRMT dataset. Acoustic data was therefore collected within and near SSP2 during the summer of 2024. This write-up is not a Biological Assessment (BA). However, the data and information provided can be used during the creation of a BA and consultations with US Fish and Wildlife Service (USFWS) in order to comply with federal directives of the Endangered Species Act of 1973 (16 U. S. C. 153 et seq.). The SSP2 site was surveyed during summer roosting/maternity season of 2024 using ultrasonic acoustic monitors to record calls from all bat species whose home ranges include the ORR. Special note was taken for presence of Federally listed Endangered and Threatened (T&E) bat species, as well as bat species which are Proposed for Federal listing, Candidate for federal listing, and state listed. Summer roosting season, from May 15 to August 15, is crucial to forest-dwelling T&E bat species for rearing young and foraging. Results of these surveys indicate the presence of three Federally listed bat species: Gray bat (Myotis grisescens--Endangered), Indiana bat (Myotis sodalis--Endangered), and Northern long-eared bat (Myotis septentrionalis--Endangered). Two additional bat species were present on the SSP2 Site: Tricolored bat (Perimyotis subflavus--Proposed for Federal listing) and Little brown bat (Myotis lucifugus—Candidate for Federal listing).
Title: Multidisciplinary Geotechnical Data Collection, Curation, and Analysis for Conformity with the Regulatory Framework for Geologic Carbon Storage in Wyoming, USA. Text: Construction and operation of wells for geologic sequestration of carbon dioxide necessitate that they are permitted under the Environmental Protection Agency’s Underground Injection Control Class VI requirements. Class VI wells conform to stringent requirements to ensure long-term safety and integrity of the storage site and the protection of Underground Sources of Drinking Water. Entities pursuing Class VI permitting must provide comprehensive geologic site characterization, including regional geologic structure and stratigraphy, aquifer information, reservoir and confining unit geomechanical properties, geochemical analyses, assessment of trapping capacity and mechanisms, and a variety of other of multidisciplinary geotechnical data. The Wyoming Class VI Site Characterization Database Project is focused on developing a geologic site characterization database of geotechnical information, which has been compiled and verified from established, public databases/entities and scientific literature to expedite Class VI permitting in Sweetwater County within the Greater Green River Basin of southern Wyoming. The preliminary suite of compiled data from 14,000 wells includes 8,000 wells with logs and 7,250 wells with formation tops, ~70 wells with core data (e.g., X-Ray diffraction, petrographic, and petrophysical data), ~2,500 water analyses, ~740 seismic events data, and ~520 bottom-hole temperature measurements. Future work on—and stemming from—this project will include new core analyses, calculation and interpolation of subsurface temperature gradients, mechanical earth models, geochemical simulations, storage capacity estimation, stratigraphic column generation and correlation, and construction of subsurface maps. Finally, this work will help to inspire and facilitate subsurface data compilation and curation beyond Sweetwater County, Wyoming.
The wider adoption of hydrogen in multiple sectors of the economy requires that safety and risk issues be rigorously investigated. Quantitative Risk Assessment (QRA) is an important tool for enabling safe deployment of hydrogen fueling stations and is increasingly embedded in the permitting process. QRA requires reliability data, and currently hydrogen QRA is limited by the lack of hydrogen specific reliability data, thereby hindering the development of necessary safety codes and standards [1]. Four tools have been identified that collect hydrogen system safety data: H2Tools Lessons Learned, Hydrogen Incidents and Accidents Database (HIAD), National Renewable Energy Lab's (NREL) Composite Data Products (CDPs), and the Center for Hydrogen Safety (CHS) Equipment and Component Failure Rate Data Submission Form. This work critically reviews and analyzes these tools for their quality and usability in QRA. It is determined that these tools lay a good foundation, however, the data collected by these tools needs improvement for use in QRA. Areas in which these tools can be improved are highlighted, and can be used to develop a path towards adequate reliability data collection for hydrogen systems.
This paper presents the findings from the ICEBERG project at Fermilab, focusing on the development and optimization of Liquid Argon Time Projection Chamber (LArTPC) detectors for the DUNE project. Various parameters, including Vref, gain, and peak time, were systematically varied to ensure accurate data collection and diagnostics. The analysis revealed optimal settings that enhance the detector's performance, paving the way for improved neutrino detection and research.
Abstract Age assessment of the living is a fundamental procedure in the process of human identification, in order to guarantee fair treatment of individuals, which has ethical, civil, legal, and medical repercussions. The careful selection of the appropriate methods requires evaluation of several parameters: accuracy, precision of the method, as well as its reproducibility. The approach proposed by Mincer et al. adapted from Demirjian et al. exploring third molar mineralisation, is one of the most frequently considered for age estimation of the living. Thus, this work aims to assess potential bias in the data collection when applying the classification stages for dental mineralisation adapted by Mincer et al. A total of 102 orthopantomographs, of clinical origin, belonging to individuals aged between 12 and 25 years ($ \bar{\textit x} $ = 20.12 years, SD = 3.49 years; 65 females, 37 males, all of Portuguese nationality) were included and a retrospective analysis performed by five observers with different levels of experience (high, average, and basic). The performance and agreement between five observers were evaluated using Weighted Cohen’s Kappa and the Intraclass Correlation Coefficient. To access the influence of impaction on third molar classification, variables were tested using ordinal logistic regression Generalised Linear Model. It was observed that there were variations in the number of teeth identified among the observers, but the agreement levels ranged from moderate to substantial (0.4–0.8). Upon closer examination of the results, it was observed that although there were discernible differences between highly experienced observers and those with less experience, the gap was not as significant as initially hypothesised, and a greater disparity between the classifications of the upper (0.24–0.49) and lower third molars (>0.55) was observed. When bone superimposition is present, the classification process is not significantly influenced; however, variation in teeth angulation affects the assessment. The results suggest that with an efficient preparation, the level of experience as a factor can be overcome. Mincer and colleague's classification system can be replicated with ease and consistency, even though the classification of upper and lower third molars presents distinct challenges.
Storm-snow avalanches are challenging to forecast due to complex alpine terrain and during rapidly changing weather conditions. They can result in loss of lives and significant economic impact. We describe how a new device that continuously measures with high-frequency snowflake mass, size, density, and type, the Differential Emissivity Imaging Disdrometer (DEID), and show how the DEID can be used to aid avalanche forecasting when coupled with a storm-snow stability model. DEID measurements of snow accumulation, snow water equivalent (SWE), and snow density obtained during seventeen storms taken at the mid-Collins Snow-Study Plot at Alta Ski Area in Utah's Central Wasatch mountain range during winter 2020–2021 show excellent agreement with infrequent manual measurements. Additionally, two new variables, the Shape Density Index (SDI) and Complexity, are proposed and used to classify snowflake habit and estimate storm-snow shear strength. We illustrate how these DEID-derived data can be used to identify layers of concern in the storm snow such as density inversions, in real-time without digging snow pits. Furthermore, the DEID-data are used to run four variations of the SNOw Slope Stability model (SNOSS) for the storms investigated. The results are evaluated with data collected from tilt-board tests, infrasound measurements, and visual observations of avalanches. For a total fourteen storms analyzed, the DEID-driven SNOSS-modeled minimum stability index predicts the general stability of the storm-snow as indicated by observed avalanches, both natural and of unknown cause. Finally, the results provide a promising approach for nowcasting instabilities within storm-snow layers with a single instrument.
Abstract. The main goal of the TRacking Aerosol Convection interactions ExpeRiment (TRACER) project was to further understand the role that regional circulations and aerosol loading play in the convective cloud life cycle across the greater Houston, Texas, area. To accomplish this goal, the United States Department of Energy and research partners collaborated to deploy atmospheric observing systems across the region. Cloud and precipitation radars, radiosondes, and air quality sensors captured atmospheric and cloud characteristics. A dense lower-atmospheric dataset was developed using ground-based remote sensors, a tethersonde, and uncrewed aerial systems (UASs). TRACER-UAS is a subproject that deployed two UAS platforms to gather high-resolution observations in the lower atmosphere between 1 June and 30 September 2022. The University of Oklahoma CopterSonde and the University of Colorado Boulder RAAVEN (Robust Autonomous Aerial Vehicle – Endurant Nimble) were flown at two coastal locations between the Gulf of Mexico and Houston. The University of Colorado Boulder RAAVEN gathered measurements of atmospheric thermodynamic state, winds and turbulence, and aerosol size distribution. Meanwhile, the University of Oklahoma CopterSonde system operated on a regular basis to resolve the vertical structure of the thermodynamic and kinematic state. Together, a complementary dataset of over 200 flight hours across 61 d was generated, and data from each platform proved to be in strong agreement. In this paper, the platforms and respective data collection and processing are described. The dataset described herein provides information on boundary layer evolution, the sea breeze circulation, conditions prior to and nearby deep convection, and the vertical structure and evolution of aerosols. The quality-controlled TRACER-UAS observations from the CopterSonde and RAAVEN can be found at https://doi.org/10.5439/1969004 (Lappin, 2023) and https://doi.org/10.5439/1985470 (de Boer, 2023), respectively.