Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “research data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

The Silencing of U.S. Campuses Following the COVID-19 Response: Evaluating Root Mean Square Seismic Amplitudes Using Power Spectral Density Data

In response to the COVID-19 global pandemic, many populated and active regions have become deserted and show significant reductions in their background seismicity, especially campuses across the United States (U.S.). Seismic sensors located in the vicinity of or within U.S. campuses show that anthropogenic seismic noise remains elevated during the ordinary, nonpandemic, academic year, only subduing during periods of recess (e.g., winter break). Here, we use power spectral density (PSD) data computed by the Incorporated Research Institutions for Seismology Data Management Center for quality assessment to calculate root mean square (rms) amplitude and analyze the effects of the COVID-19 school closures. We processed and analyzed PSD data for 46 seismic stations located within 50 m of a U.S. university or college. Results show that 42 campus stations show an overall rms drop following a statewide school closure.

58 GEOSCIENCES↗

From Data to Discovery: AI's Transformative Role in Thin Film Research

The advancement of thin film technologies is pivotal for progress in numerous fields, including energy, electronics, and quantum computing. However, the traditional trial-and-error approach to materials discovery is inherently slow and inefficient. This presentation will showcase how artificial intelligence (AI) is transforming thin film research by enabling a data-driven paradigm shift. We will highlight our past successes in applying AI to understand radiation damage in thin film oxides, demonstrating how graph analytics can unravel complex material behavior. Additionally, we will provide insights into our current work at the National Renewable Energy Laboratory, where we are leading the charge in autonomous materials science. Backed by a $14M investment in our characterization facility, we are developing AI-guided workflows that seamlessly integrate experimentation and AI-guided decision-making. By harnessing the power of AI, we aim to accelerate the discovery and design of high-performance thin films, propelling innovation across a multitude of industries.

36 MATERIALS SCIENCE↗

Enhancing Discoverability and Management of Atmospheric Data at Scale: Solutions from the ARM Data Center

The Atmospheric Radiation Measurement (ARM) is a multi-laboratory and multi-institutional U.S. Department of Energy (DOE) Office of Science National User Facility. The ARM Data Center (ADC), located at Oak Ridge National Laboratory, collects, archives, and shares vast atmospheric data crucial for climate research. The ADC manages over 7 PB of data from 460 instruments worldwide, processing it into more than 11,000 diverse data products using the Network Common Data Form (NetCDF) for machine-independent accessibility. The primary challenge addressed in this paper is the efficient management and distribution of vast and diverse datasets essential for the climate research community, enhancing accessibility through advanced tools like Data Discovery. The ADC has developed advanced infrastructure and software architecture to handle the continuous influx of heterogeneous data to enhance data discoverability, resulting in increased scientific collaboration. In 2023, users from over 34 countries downloaded and utilized ARM data, resulting in 1,455 publications. The ADC’s efforts have significantly improved the discoverability and usability of atmospheric data, fostering extensive scientific research and collaboration. This paper details the solutions implemented by the ADC team for efficient data discovery and distribution, and it demonstrates ARM’s capability of staging processed data for scientific analysis.

Shah, Chirag [ORNL] (ORCID:0000000203145737)↗

Dataset for 'Ombadi, M. & Varadharajan, C. (2022). Urbanization and aridity mediate distinct salinity response to floods in rivers and streams across the Contiguous United States, Water Research'

This package contains data sets and code used to obtain the results in Ombadi, M., & Varadharajan, C. (2022). Urbanization and aridity mediate distinct salinity response to floods in rivers and streams across the Contiguous United States. Water Research, 118664. The folder "data" contains 259 .csv files, each of which has daily time series of concurrent streamflow (Q) and specific conductance (SC) for each of the sites used in this study originally downloaded from the USGS National Water Information System (NWIS; USGS, 2016). The number of data points in each of the files is at least 3650 (i.e. 10 years of daily measurements). The folder "RF_single_data" contains 259 .csv files, each of which include data used to train and test the Random Forest models at individual sites for predicting SC during days of floods. The folder "RF_regional_data" contains 3 .csv files, each of which include scaled data compiled from all sites within each climate zone (arid, temperate and wet). "metadata.csv" contains the physical properties of the 259 catchments corresponding to the sites used in this study; this data was extracted from GAGES-II dataset (Falcone et al., 2010). "RF_implementation.ipynb" is a Jupyter notebook with the code needed to implement the analysis using Random Forest models either for individual sites or for the regional models (for each climate zone). The code utilizes the data in the two folders: "RF_single_data" and "RF_regional_data" and the metadata.csv file.

54 ENVIRONMENTAL SCIENCES↗

The Coastal Carbon Library and Atlas: Open source soil data and tools supporting blue carbon research and policy

Abstract Quantifying carbon fluxes into and out of coastal soils is critical to meeting greenhouse gas reduction and coastal resiliency goals. Numerous ‘blue carbon’ studies have generated, or benefitted from, synthetic datasets. However, the community those efforts inspired does not have a centralized, standardized database of disaggregated data used to estimate carbon stocks and fluxes. In this paper, we describe a data structure designed to standardize data reporting, maximize reuse, and maintain a chain of credit from synthesis to original source. We introduce version 1.0.0. of the Coastal Carbon Library, a global database of 6723 soil profiles representing blue carbon‐storing systems including marshes, mangroves, tidal freshwater forests, and seagrasses. We also present the Coastal Carbon Atlas, an R‐shiny application that can be used to visualize, query, and download portions of the Coastal Carbon Library. The majority (4815) of entries in the database can be used for carbon stock assessments without the need for interpolating missing soil variables, 533 are available for estimating carbon burial rate, and 326 are useful for fitting dynamic soil formation models. Organic matter density significantly varied by habitat with tidal freshwater forests having the highest density, and seagrasses having the lowest. Future work could involve expansion of the synthesis to include more deep stock assessments, increasing the representation of data outside of the U.S., and increasing the amount of data available for mangroves and seagrasses, especially carbon burial rate data. We present proposed best practices for blue carbon data including an emphasis on disaggregation, data publication, dataset documentation, and use of standardized vocabulary and templates whenever appropriate. To conclude, the Coastal Carbon Library and Atlas serve as a general example of a grassroots F.A.I.R. (Findable, Accessible, Interoperable, and Reusable) data effort demonstrating how data producers can coordinate to develop tools relevant to policy and decision‐making.

Holmquist, James R.↗

Multimodal Chemical Characterization of Brown Carbon in Atmosphere and Snowpack during SAIL Field Campaign Report

The main goal of this field project was to characterize the chemical and physical properties of light-absorbing particles (LAP) present in the atmosphere and snowpack. The specific field efforts included: 1) online measurement and sampling of light-absorbing particles using a Magee Scientific model AE33 aethalometer (AE33) and a time-resolved aerosol collector (TRAC); and 2) sampling of snowpack during observed events of LAP deposition on snow, confirmed by onsite U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility measurements and meteorological data. The ongoing research objectives of our follow-up sample and data analysis include: 1) Characterization of aerosol regimes and sources, identifying particle-type populations, mixing states, and atmospheric transformations based on correlative analysis of our detailed chemical imaging and chemical characterization measurements and real-time records from the second ARM Mobile Facility (AMF2) and other instruments available from the Surface Atmosphere Integrated Field Laboratory (SAIL) experiment; 2) Assessment of the optical and chemical properties of snowpack deposits to investigate how aerosol deposition influences snowpack lifetime.

54 ENVIRONMENTAL SCIENCES↗

TSDC: Transportation Secure Data Center: Real-World Data for Planning, Modeling, and Analysis

The Transportation Secure Data Center is a centralized repository for high-resolution transportation data from hundreds of travel and transit surveys and studies. It makes vital transportation data broadly available to users while preserving the privacy of survey participants. It houses surveys and studies conducted by state departments of transportation, metropolitan planning organizations, transit agencies, cities, and other public agencies. Meanwhile, the Livewire Data Platform empowers research, industry, and academic partners to easily and securely preserve, maintain, share, discover, and gain access to transportation and mobility data. Livewire accommodates a range of datasets, including behavioral, experimental, model, analytical, and raw data at the vehicle, traveler, and system levels. Datasets support mobility research and planning spanning urban science, connected and automated vehicles, fueling and charging infrastructure, mobility decision science, multimodal transportation, vehicle efficiency, and more.

33 ADVANCED PROPULSION SYSTEMS↗

Impact of Domain Knowledge on the Property Prediction of Specialized Machine Learning Models

Developing transferable machine learning models is trending in data-driven materials research. However, how to apply such models to a specific research domain remains unclear. Here, in this work, we choose high-entropy materials as a platform with a specialized data set containing 145,323 DFT-relaxed materials. This data set is used to explore the role of domain-specific knowledge in training effective models. Our tests with three representative graph neural network architectures indicate the model complexity has much smaller influence on performance than the data itself. Specifically, the consideration of low-energy atomic ordering, structures with diverse elemental coverage, and high-order interactions significantly influences the model performance. We also find that domain knowledge-driven sampling can greatly enhance unsupervised learning techniques. This research highlights that developing specialized data sets is more beneficial than further complicating deep learning architectures. Additionally, physics-inspired sampling algorithms are crucially needed for better machine learning models for a specific materials research domain.

36 MATERIALS SCIENCE↗

Time for a drought experiment: Do you know your plants’ water status?

Abstract Drought stress is an increasing concern because of climate change and increasing demands on water for agriculture. There are still many unknowns about how plants sense and respond to water limitation, including which genes and cellular mechanisms are impactful for ecology and crop improvement in drought-prone environments. A better understanding of plant drought resistance will require integration of several research disciplines. A common set of parameters to describe plant water status and quantify drought severity can enhance data interpretation and research integration across the research disciplines involved in understanding drought resistance and would be especially useful in integrating the flood of genomic data being generated in drought studies. Water potential (ψw) is a physical measure of the free energy status of water that, along with related physiological measurements, allows unambiguous description of plant water status that can apply across various soil types and environmental conditions. ψw and related physiological parameters can be measured with relatively modest investment in equipment and effort. Thus, we propose that increased use of ψw as a fundamental descriptor of plant water status can enhance the insight gained from many drought-related experiments and facilitate data integration and sharing across laboratories and research disciplines.

Juenger, Thomas E. (ORCID:0000000195509288)↗

The disCO2ver Platform: Curating Data and Tools for Geologic Carbon Sequestration and Deep Subsurface Research Systems

The U.S. DOE National Energy Technology Laboratory has invested 12+ years of development into the data repository and digital laboratory, the Energy Data eXchange (EDX, edx.netl.doe.gov). Supporting a variety of research areas across the DOE Office of Fossil Energy and Carbon Management, the platform has successfully curated and preserved thousands of data products from DOE research. The Carbon Storage Program has successfully supported data curation, upload, and publishing of data products on EDX for many years, demonstrating a success story of how resources like EDX can effectively help with long term preservation and publishing of DOE data products. EDX continues to shift towards cloud-supported infrastructure, taking a hybrid approach combining on-premises compute and storage integrated with cloud-hosted services. The integration of cloud compute and hybrid architecture enables the development of EDX-hosted platforms that tailor the data and tools hosted on them to a specific community, enables implementation of machine learning tools for data discovery and filtering, and enables the hosting of virtual (online user interface) tools. Geologic carbon sequestration (GCS) research continues to scale up in response to the current administration goals to reduce greenhouse gas emissions and transition the energy economy. Over the last year, EDX’s disCO2ver platform has been developed in response to the need for access to data products and tools to support the scaling up of GCS research. disCO2ver provides access to data resources and tools, produced by DOE and outside authoritative external resources. The platform also provides a user-access control component for the virtualization and cloud hosting of tools. Tools that need to be virtualized, to eliminate the need for users to download the tool and use local compute resources, is essential to supporting big-data analysis and machine learning that is becoming common place in carbon storage modeling, risk analysis, and data publishing practices. This talk will review the EDX’s disCO2ver platform and the current work ongoing to curate data and tools to support GCS and deep subsurface systems research.

Morkner, Paige↗

Breaking the barrier of human-annotated training data for machine learning-aided plant research using aerial imagery

Machine learning (ML) can accelerate biological research. However, the adoption of such tools to facilitate phenotyping based on sensor data has been limited by (i) the need for a large amount of human-annotated training data for each context in which the tool is used and (ii) phenotypes varying across contexts defined in terms of genetics and environment. This is a major bottleneck because acquiring training data is generally costly and time-consuming. This study demonstrates how a ML approach can address these challenges by minimizing the amount of human supervision needed for tool building. A case study was performed to compare ML approaches that examine images collected by an uncrewed aerial vehicle to determine the presence/absence of panicles (i.e. “heading”) across thousands of field plots containing genetically diverse breeding populations of 2 Miscanthus species. Automated analysis of aerial imagery enabled the identification of heading approximately 9 times faster than in-field visual inspection by humans. Leveraging an Efficiently Supervised Generative Adversarial Network (ESGAN) learning strategy reduced the requirement for human-annotated data by 1 to 2 orders of magnitude compared to traditional, fully supervised learning approaches. The ESGAN model learned the salient features of the data set by using thousands of unlabeled images to inform the discriminative ability of a classifier so that it required minimal human-labeled training data. This method can accelerate the phenotyping of heading date as a measure of flowering time in Miscanthus across diverse contexts (e.g. in multistate trials) and opens avenues to promote the broad adoption of ML tools.

59 BASIC BIOLOGICAL SCIENCES↗

Accelerating Low-Income Financing and Transactions (LIFT) for Solar Access Everywhere (Final Technical Report)

The Accelerating Low-Income Financing and Transactions (LIFT) for Solar Access Everywhere project’s goal was to expand Low-to-Moderate Income (LMI) solar access for homeowners and renters. The LIFT project researched and gathered data on 453 LMI community solar project across the country. Following three years of research, the project delivered three groundbreaking research papers in June 2022, focused on 1) customer experience, 2) the growth of community solar programs, and 3) project-level financial best practices for serving LMI communities. These were followed by a user-friendly web-based Toolkit allowing users to interact with project data and key findings in November 2022. The customer experience research examined community solar subscribers’ primary motivations to join and remain satisfied with projects. Our research identified 453 projects across the country that dedicated some portion of the system capacity to LMI households. Seventeen of these projects participated in the LIFT customer experience research, allowing the project team to survey their customers and gain insight into how LMI subscribers feel about community solar and the programs that serve them. Subscribers in our sample indicated that the most critical issue that motivated them to participate in their program, however, was not savings but helping the environment. This was true for both LMI and non-LMI subscribers. Helping the environment was also the most important issue for LMI subscribers to measure how well their program was working for them. LIFT also explored how rapidly community solar has grown since its inception in 2006, publishing results in the Growth of U.S. Community Solar Serving LMI Households report. The results showed that community solar projects serving LMI households are one of the fastest growing segments of the solar industry. The report identifies and recommends ways developers should overcome real or perceived risks to LMI customer acquisition and subscriber management. Through the analysis of community solar project finance research, LIFT showed that most community solar projects serving LMI households are financed in the same ways mainstream community solar projects are financed. The value stacks and financial returns are no different, although LMI inclusion and participation rate varied across programs in our sample, ranging from between 10% and 100%. Based on the findings from the LIFT research, the team built a web-based user-friendly Toolkit, consisting of case studies, project finance best practices, and several tools built around the national dataset of 453 community solar projects that serve LMI households. These allow users to engage with the dataset in multiple ways; to explore the landscape of LMI community solar in the U.S., and to design community solar projects to optimize LMI inclusion, equity, and savings levels. The Toolkit also includes a library of LIFT-generated and LIFT-curated resources for users to learn more about how to best serve LMI communities through community solar. LIFT officially published the Toolkit on October 31, 2022, followed by a launch event (public webinar) on November 17, 2022. The core LIFT partners continue to engage in outreach and dissemination efforts to promote the LIFT Toolkit and research publications. Our driving motivation is to continue enabling solar developers to leverage the findings of this three-year research effort. By implication, the LIFT Toolkit is designed for use by utilities, energy service providers, and financiers or investors as a learning and decision-making tool to rapidly scale project models that optimize LMI inclusion and maximize real household savings.

14 SOLAR ENERGY↗

Examining United States Nuclear Weapon Educational Initiatives_A Quantitative Research Study on the Satisfaction of Participants Within the Nuclear Weapons Community

This study focused on a few United States nuclear weapons educational initiatives exploring six classroom environment factors that may or may not influence Satisfaction in a virtual or live classroom environment within the nuclear weapons community. The research explored Personalization, Involvement, Student Cohesiveness, Task Orientation, Innovation, and Individualization. The United States nuclear weapons community needed to attract new and talented individuals to the nuclear weapons profession. This study added to the nuclear weapons community by researching factors that may aid academic instructors, and curriculum developers create new and inventive ways of incorporating these factors into existing or new nuclear weapons educational courses. Convenience sampling was the data collection method used to gather quantitative research data. One hundred and seven voluntary participants utilized the College and University Classroom Environment Inventory with the Los Alamos National Laboratory's consent. There is not much concurrent research on Satisfaction within the live and virtual classroom environment. This research was vital for future research into how to increase participant satisfaction. The study provided the nuclear weapons community with grounded analysis to help create and guide new classes and course creation, possibly increasing Satisfaction in a live and virtual classroom environment. According to binary logistic regression, there was a statistically significant relationship among the dependent variable, Satisfaction, and two independent variables, Involvement, and Task Orientation in the classroom environment within the nuclear weapons community. Future nuclear educational programs may increase Satisfaction in the classroom environment by increasing participant Involvement and Task Orientation within a class using the CUCEI as a reference.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Time to Use Dendrohydrological Data in Water Resources Management?

Despite a wide consensus on the potential importance of streamflow reconstructions to water resources planning and management activities, applications of dendrohydrology have been piecemeal. This review article describes barriers to widespread use of dendrohydrological data and suggests research needs for operationalizing streamflow records reconstructed from these proxy data.

42 ENGINEERING↗

RCSB Protein Data Bank tools for 3D structure-guided cancer research: human papillomavirus (HPV) case study

Abstract Atomic-level three-dimensional (3D) structure data for biological macromolecules often prove critical to dissecting and understanding the precise mechanisms of action of cancer-related proteins and their diverse roles in oncogenic transformation, proliferation, and metastasis. They are also used extensively to identify potentially druggable targets and facilitate discovery and development of both small-molecule and biologic drugs that are today benefiting individuals diagnosed with cancer around the world. 3D structures of biomolecules (including proteins, DNA, RNA, and their complexes with one another, drugs, and other small molecules) are freely distributed by the open-access Protein Data Bank (PDB). This global data repository is used by millions of scientists and educators working in the areas of drug discovery, vaccine design, and biomedical and biotechnology research. The US Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) provides an integrated portal to the PDB archive that streamlines access for millions of worldwide PDB data consumers worldwide. Herein, we review online resources made available free of charge by the RCSB PDB to basic and applied researchers, healthcare providers, educators and their students, patients and their families, and the curious public. We exemplify the value of understanding cancer-related proteins in 3D with a case study focused on human papillomavirus.

60 APPLIED LIFE SCIENCES↗

Black-box statistical prediction of lossy compression ratios for scientific data

Lossy compressors are increasingly adopted in scientific research, tackling volumes of data from experiments or parallel numerical simulations and facilitating data storage and movement. In contrast with the notion of entropy in lossless compression, no theoretical or data-based quantification of lossy compressibility exists for scientific data. Users rely on trial and error to assess lossy compression performance. As a strong data-driven effort toward quantifying lossy compressibility of scientific datasets, we provide a statistical framework to predict compression ratios of lossy compressors. Our method is a two-step framework where (i) compressor-agnostic predictors are computed and (ii) statistical prediction models relying on these predictors are trained on observed compression ratios. Proposed predictors exploit spatial correlations and notions of entropy and lossyness via the quantized entropy. We study 8+ compressors on 6 scientific datasets and achieve a median percentage prediction error less than 12%, which is substantially smaller than that of other methods while achieving at least a 8.8× speedup for searching for a specific compression ratio and 7.8× speedup for determining the best compressor out of a collection.

97 MATHEMATICS AND COMPUTING↗

Summary Report of the 2nd RCM of the CRP on Updating Fission Yield Data for Applications

The Second Research Coordination Meeting of the IAEA Coordinated Research Project (CRP) on Updating Fission Yield Data for Applications was held in Vienna at the IAEA headquarters from 19 to 23 December 2022, with 23 international experts attending the meeting. The CRP is devoted to evaluation efforts of cumulative and independent fission yields for incident energies from the thermal point up to 14 MeV on actinide targets. Produced fission yield evaluations should include full uncertainty quantification and are expected to combine available experimental data and state-of-the-art model information. The activities undertaken within this CRP were reviewed including the assessment of newly measured data and ongoing evaluation efforts. Technical discussions and the resulting further work plan of this CRP are summarized in this report. The meeting presentations are available at: https://www-nds.iaea.org/index-meeting-crp/2RCM_FY/index.htm.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗