Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

BLDAP Intro to Python/Data Science Curriculum v1

The Github repository contains the Jupyter notebooks for the intro to Python / Data Science course for Berkeley Lab Director's Apprenticeship Program (BLDAP). This course is designed for students with little to no experience in coding to learn skills in Python necessary for data science. Students utilize Jupyter notebooks throughout the course. The overall goal is for students to learn how to use Python to clean, analyze, and visualize large data sets in order to communicate effectively their conclusions about the data set. Students apply the skills they learned on actual data sets provided by researchers in Berkeley Lab.

Hales, Laurel [Lawrence Berkeley National Laborato↗

An open-access simulated earthquake ground-motion database for an M7 Hayward Fault earthquake in the San Francisco Bay Region

Comprehensive understanding of earthquake ground motions, particularly in the near-fault region of large-magnitude events, is limited by gaps in strong-motion data. This challenge is prominent in areas with high seismic hazard but infrequent large earthquakes where data is sparse and difficult to interpret. These data limitations lead to uncertainties in the development of site-specific ground motions, which are crucial for engineering risk assessments. To address these challenges, physics-based regional-scale ground-motion simulations have been developed. With the emergence of exaflop-scale computing ecosystems, it is now possible to simulate regional earthquake processes at unprecedented fidelity and generate the large number of fault rupture realizations necessary to characterize both intra- and inter-event ground-motion variability. This article introduces a new database of simulated earthquake ground motions, created for applications in earthquake engineering, earthquake planning, and emergency response. The inaugural version of the database features simulated ground motions for a magnitude 7 Hayward Fault earthquake in the San Francisco Bay Region (SFBR), using the EarthQuake SIMulation (EQSIM) simulation framework and the Graves–Pitarka kinematic rupture model. The aim is to provide high-fidelity, spatially dense, three-component motions generated on the Department of Energy’s (DOE) newest generation of graphics processing unit (GPU)-accelerated supercomputers. These motions are being made openly available to the engineering, scientific, and disaster planning communities. In addition, this work develops protocols for the efficient dissemination of these large data sets and emphasizes community engagement to build confidence in their application. This article discusses the methodology behind the data, underlying software verification and validation, scalable data management, and a user interface for data access. The goal is to facilitate widespread use and elicit expert feedback to maximize the utility and exploitation of simulated motions. While the initial focus is on the San Francisco Region, simulations for additional regions will be added as the DOE program progresses.

Simulated ground-motion database↗

Automatic Waveform Quality Control for Surface Waves Using Machine Learning

Surface-wave seismograms are widely used by researchers to study Earth’s interior and earthquakes. To extract information reliably and robustly from a suite of surface waveforms, the signals require quality control screening to reduce artifacts from signal complexity and noise. This process has usually been completed by human experts labeling each waveform visually, which is time consuming and tedious for large data sets. We explore automated approaches to improve the efficiency of waveform quality control processing by investigating logistic regression, support vector machines, K-nearest neighbors, random forests (RF), and artificial neural networks (ANN) algorithms. To speed up signal quality assessment, we trained these five machine learning (ML) methods using nearly 400,000 human-labeled waveforms. The ANN and RF models outperformed other algorithms and achieved a test accuracy of 92%. We evaluated these two best-performing models using seismic events from geographic regions not used for training. The results show that the two trained models agree with labels from human analysts but required only 0.4% of the time. Although the original (human) quality assignments assessed general waveform signal-to-noise, the ANN or RF labels can help facilitate detailed waveform analysis. Our investigations demonstrate the capability of the automated processing using these two ML models to reduce outliers in surface-wave-related measurements without human quality control screening.

58 GEOSCIENCES↗

TEACHING AN OLD ACCELERATOR NEW TRICKS

The Argonne Tandem Linac Accelerator System (ATLAS) has been a National User Facility since 1985. In that time, many of the systems that help operators retrieve, modify, and store beamline parameters have not kept pace with the advancement of technology. Development of a new method of storing and retrieving beamline parameters resulted in the testing and installation of a time-series database as a potential replacement for the traditional relational database. InfluxDB was selected due to its self-hosted Open-Source version availability as well as the simplicity of installation and setup. A program was written to periodically gather all accelerator parameters in the control system and store them in the time-series database. This resulted in over 13,000 distinct data points, captured at 5-minute intervals. A second test captured 35 channels on a 1-minute cadence. Graphing of the captured data is being done on Grafana, an Open-Source version is available that co-exists well with InfluxDB as the back-end. Grafana made visualizing the data simple and flexible. The testing has allowed for the use of modern graphing tools to generate new insights into operating the accelerator, as well as opened the door to building large data sets suitable for Artificial Intelligence and Machine Learning applications.

Novak, D.↗

Water resource recovery modelling 2021 (WRRmod2021 conference)

Our society is transitioning fast into the digital age, spurred by development of cheap and new sensing technology, breakthroughs in computing, and development of efficient algorithms for optimization. This transition is also visible in the field of wastewater treatment and is driving new model developments, especially by exploiting the large data sets available with many utilities. Not surprisingly, WRRmod2021 had featured a strong session on ‘data-driven models and digitalization’ focused on this hot topic. At the same time, engineering practice calls for more robust models for performance evaluation and optimization of both conventional facilities and innovative processes. As a result, the WRRmod2021 program also exhibited sessions on modelling of new process units (e.g., aerobic granular sludge), modelling of the nitrogen cycle, and integrated/plant-wide modelling.

54 ENVIRONMENTAL SCIENCES↗

Nuclear Physics Network Requirements Review Report

The Energy Sciences Network (ESnet) is the Office of Science’s high-performance network user facility, delivering highly reliable data transport capabilities optimized for the requirements of data-intensive science. In essence, ESnet is the circulatory system that enables the U.S. Department of Energy (DOE) science mission by connecting each and every DOE lab and its user facilities. ESnet is funded and stewarded by the Advanced Scientific Computing Research (ASCR) Program and managed and operated by the Scientific Networking Division at Lawrence Berkeley National Laboratory (LBNL). ESnet is widely regarded as a global leader in the research and education networking community. ESnet connects DOE national laboratories, user facilities, and major experiments so scientists can use remote instruments and computing resources as well as share data with collaborators, transfer large data sets, and access distributed data repositories. While ESnet provides network connectivity, it cannot be characterized as an internet service provider as it is specifically built to provide a range of network services that are tailored to meet the unique requirements of DOE’s data-intensive science.

97 MATHEMATICS AND COMPUTING↗

Redox transistors based on TiO 2 for analogue neuromorphic computing

The ability to train deep neural networks on large data sets have made significant impacts onto artificial intelligence, but consume significant amounts of energy due to the need to move information from memory to logic units. In-memory "neuromorphic" computing presents an alternative framework that processes information directly on memory elements. In-memory computing has been limited by the poor performance of the analogue information storage element, often phase-change memory or memristors. To solve this problem, we developed two types of "redox transistors" using TiO 2 (anatase) which stores analogue information states through the electrochemical concentration of dopants in the crystal. The first type of redox transistor uses lithium as the electrochemical dopant ion, and its key advantage is low operating voltage. The second uses oxygen vacancies as the dopant, which is CMOS compatible and can retain state even when scaled to nanosized dimensions. Both devices offer significant advantages in terms of predictable analogue switching over conventional filamentary-based devices, and provide a significant advance in developing materials and devices for neuromorphic computing.

36 MATERIALS SCIENCE↗

Identifying Suspect Instrument Intervals Using Midnight Noise Time Histories

With an intense work ethic, and high levels of devotion to their craft, seismologists continue to battle the elements and wrestle with complex logistics in their efforts to field instrumentation over the entire globe. We are fortunate to see data volumes accelerating in size, and quality continuing to improve; for example, universal timing issues can now be considered rare. However, instrument response issues are still common, and can hinder research that seeks to understand amplitudes of seismic wavefields. We are interested in using the accumulations of seismic data to develop models that will predict high frequency (0.2-20 Hz, and higher) signal amplitudes for explosion monitoring purposes over broad areas, which requires large data sets and extensive quality control. To identify response and station health issues, we have collected noise time histories for global seismic data, focusing on measurements near (but not restricted to) midnight to eliminate diurnal variations, and have manually determined time intervals that appear inconsistent with background behavior. We assign descriptive labels, but do not attempt to diagnose causes. We use these intervals to discard data. To date, we have examined 39,260 channels from 11,105 stations, heavily weighted toward IRIS holdings, dates through 2017 (depending on station, roughly the date we started our manual review), bands between 1 and 8 Hz, finding 24,733 anomalous time intervals. The great majority (90%) of these intervals appear to be shifts of constant offset, often bounded by times of known instrument changes, likely the result of poor documentation of response parameters at one of many stages between the field and the plotter. We hope these results can be of use to our colleagues, and would encourage community efforts to diagnose anomalous behavior, and fix poor responses. We also hope these results will support automation efforts, including application of supervised learning techniques.

47 OTHER INSTRUMENTATION↗

Use of the Tool to Support Renewable Energy Auctions Processes (RE Data Explorer)

Renewable energy auctions are now a common competitive approach to procure low-cost renewable power around the world. Ensuring a successful auction process increasingly depends on the capabilities of auction designers and participants to identify actionable and defensible insights from large data sets (on renewable energy resources and complementary data) to both attract potential investors and address stakeholder concerns. The Renewable Energy (RE) Data Explorer is a user-friendly geospatial analysis tool for analyzing renewable energy potential and informing decisions. Developed by the National Renewable Energy Laboratory (NREL) and supported by the U.S. Agency for International Development (USAID), RE Data Explorer performs visualization and analysis of renewable energy potential that can be customized for different scenarios. RE Data Explorer can support prospecting, integrated planning, policymaking, and other decision-making activities to accelerate renewable energy deployment. The broader RE Explorer website provides guidance and information to link the RE Data Explorer geospatial analysis tool to key decision areas. This document provides information on how the RE Data Explorer can be used to support renewable energy auction processes.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Rheo-Structural Spectroscopy: Fingerprinting the In Situ Response of Fluids to Arbitrary Flow Fields

The objectives of this project were to develop new sample environments, measurement methodologies and associated modeling tools for characterizing the structural response to arbitrarily complex processing flows using small angle scattering, and to apply these new tools for understanding the fundamental physics governing the structuring of anisotropic particulate and polymeric materials under flow histories and conditions relevant to industrial processing flows. The research resulted in the development and implementation of a new sample environment, the fluidic four roll mill (FFoRM), for in situ small angle neutron and X-ray scattering (SANS/SAXS) measurements. These measurements are capable of generating large data sets that “fingerprint” how a complex fluid responds to a wide range of flow histories involving time variations in deformation type and rate. New modeling tools were developed to extract detailed microstructural information from such data sets, including orientation distribution functions and interparticle correlation functions, as well as reduced-order parametric descriptors of these high-dimensional functions that can be used to readily map, visualize and interpret a fluid’s structural response to its flow history. These new tools were applied to a range of model materials involving elongated particle suspensions in order to provide new insights into the physics of how flow couples with orientational and structural order in complex flows, particularly under non-dilute conditions for which no accurate theories currently exist. Using these investigations, we elucidated a number of new insights into the fundamental phenomena driving such process-structure-property relationships. These findings provide guidance for the further development of rheological models, and ultimately can inform the rational and model-based design of flow processes to achieve optimized orientational ordering that is key to the properties and function of a wide range of energy-relevant materials.

36 MATERIALS SCIENCE↗

Simultaneous mitigation of density and energy errors in approximate DFT for transition metal chemistry (Final Technical Report)

There were three major goals and objectives of this project: 1) Develop tools to understand density-driven errors in transition metal complexes, 2) Evaluate and minimize energy delocalization error and static correlation error through judicious functional choice, and 3) applying this workflow to machine-learning accelerated screening of redox couples. Over the reporting period, we developed a framework for eliminating flat plane errors. We introduced fully non-empirical coefficients. We demonstrated the approach on both molecules and solids. We investigated and eliminated density driven errors and demonstrated their impact on potential energy surfaces. We trained machine learning models both in a method-dependent fashion and to predict errors in method accuracy. We built large data sets of small molecule energetics and multi-reference character.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Data Science and Machine Learning in Education

The growing role of data science (DS) and machine learning (ML) in high-energy physics (HEP) is well established and pertinent given the complex detectors, large data, sets and sophisticated analyses at the heart of HEP research. Moreover, exploiting symmetries inherent in physics data have inspired physics-informed ML as a vibrant sub-field of computer science research. HEP researchers benefit greatly from materials widely available materials for use in education, training and workforce development. They are also contributing to these materials and providing software to DS/ML-related fields. Increasingly, physics departments are offering courses at the intersection of DS, ML and physics, often using curricula developed by HEP researchers and involving open software and data used in HEP. In this white paper, we explore synergies between HEP research and DS/ML education, discuss opportunities and challenges at this intersection, and propose community activities that will be mutually beneficial.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

IDAES Enterprise: Generation Expansion Planning with Enhanced Requirements for Capacity Adequacy Under Renewable Intermittency

Achieving net zero carbon emissions likely requires future power systems to integrate new, flexible energy technologies to accommodate higher levels of capacity from variable renewable energy sources. To determine the optimal deployment of new electricity capacity and to study the likelihood of deployments of new energy technologies, an expansion planning model has been developed as part of the IDAES-Enterprise suite of grid models. The Generation Expansion Planning (GEP) model is a multi-period model in which investment decisions occur yearly, and a Unit Commitment (UC) problem is examined on an hourly timescale. To reduce computational complexity of the GEP model, the UC problem is solved for average “representative days” which leaves out extreme, but relatively common, scenarios in which low renewable generation occurs, leaving the system with inadequacy in capacity. The IDAES-Enterprise GEP model has been modified to include these extreme scenarios while keeping the model reasonably tractable. Specifically, a lazy constraint technique was implemented to check for capacity adequacy on an hourly basis over a large data set of aligned load-wind-solar profiles. As a vast majority of the capacity constraints will not be violated, the technique lowers computational expense by searching for violated capacity constraints over an “iterative manner,” adding those infeasible constraints back into the model. Results on a test case of the Southwest Power Pool shows that the lazy constraint technique significantly reduces retirements and increases installments of natural gas combined cycles and flexible natural gas units. It also reduces some retirements of coal units. These modifications provide a more reasonable estimation of required dispatchable power generation capacity to ensure feasibility during peak net load.

Liu, Peng↗

Measurement of cross sections for mesonless charged-current muon neutrino interactions on argon with and without protons at MicroBooNE

urrent and upcoming neutrino experiments at Fermilab will rely on liquid argon time projection chambers (LArTPCs) as the primary detector technology. To reach the ambitious precision required for their physics goals, a detailed understanding of neutrino-argon scattering must be achieved. For the present Short-Baseline Neutrino program, the dominant reaction channel is charged-current muon neutrino interactions leading to mesonless final states with and without protons. Using a large data set of neutrino interactions recorded by the MicroBooNE LArTPC detector at Fermilab, we report progress towards new measurements of differential cross sections in this channel. The analysis leverages multiple event reconstruction paradigms and observables specifically chosen to maximize sensitivity to nuclear effects of interest for improving neutrino scattering models. Related previous cross-section results from MicroBooNE and a preview of what to expect from the new analysis will both be featured in this presentation.

Liu, Liang [Fermilab]↗

Gamma Ray Source Localization for Time Projection Chamber Telescopes Using Convolutional Neural Networks

Diverse phenomena such as positron annihilation in the Milky Way, merging binary neutron stars, and dark matter can be better understood by studying their gamma ray emission. Despite their importance, MeV gamma rays have been poorly explored at sensitivities that would allow for deeper insight into the nature of the gamma emitting objects. In response, a liquid argon time projection chamber (TPC) gamma ray instrument concept called GammaTPC has been proposed and promises exploration of the entire sky with a large field of view, large effective area, and high polarization sensitivity. Optimizing the pointing capability of this instrument is crucial and can be accomplished by leveraging convolutional neural networks to reconstruct electron recoil paths from Compton scattering events within the detector. In this investigation, we develop a machine learning model architecture to accommodate a large data set of high fidelity simulated electron tracks and reconstruct paths. We create two model architectures: one to predict the electron recoil track origin and one for the initial scattering direction. We find that these models predict the true origin and direction with extremely high accuracy, thereby optimizing the observatory’s estimates of the sky location of gamma ray sources.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A Comprehensive Characterization of Galaxy-cool CGM Connections at z < 0.4 with DESI Year 1 Data

We investigate the relationships between the cool circumgalactic medium (CGM), traced by Ca II absorption lines, and galaxy properties at z < 0.4 using ∼900,000 galaxy–quasar pairs within 200 kpc from the Year 1 data of the Dark Energy Spectroscopic Instrument (DESI). This large data set enables us to obtain composite spectra with sensitivity reaching to the mÅ level and to explore the Ca II absorption as a function of stellar mass, star formation rate (SFR), redshift, and galaxy types, including active galactic nuclei (AGNs). Our results show a positive correlation between the absorption strength and stellar mass of star-forming galaxies with $\langle$$W$$^{Ca II}_{0}$$\rangle$ α $M$$^{0.5}_{*}$ over 3 orders of magnitude in stellar mass from ∼10 8 to 10 11 M ⊙ , while such a mass dependence is weaker for quiescent galaxies. At a fixed mass, Ca II absorption is stronger around star-forming galaxies than quiescent ones especially within impact parameters <30 kpc. Among star-forming galaxies, the Ca II absorption further correlates with SFR, following ∝SFR 0.3 . However, in contrast to the results at higher redshifts, stronger absorption is not preferentially observed along the minor axis of star-forming galaxies, indicating a possible redshift evolution of CGM dynamics resulting from galactic feedback. Moreover, no significant difference between the properties of the cool gas around AGNs and galaxies is detected. Finally, we measure the absorption profiles with respect to the virial radius of dark matter halos and show that the total Ca II mass in the CGM is comparable to the Ca mass in the ISM of galaxies.

Circumgalactic medium↗

Vehicle Powertrain Simulation Accuracy for Various Drive Cycle Frequencies and Upsampling Techniques

As connected and automated vehicle technologies emerge and proliferate, lower frequency vehicle trajectory data is becoming more widely available. In some cases, entire fleets are streaming position, speed, and telemetry at sample rates of less than 10 seconds. This presents opportunities to apply powertrain simulators such as the National Renewable Energy Laboratory's Future Automotive Systems Technology Simulator to model how advanced powertrain technologies would perform in the real world. However, connected vehicle data tends to be available at lower temporal frequencies than the 1-10 Hz trajectories that have typically been used for powertrain simulation. Higher frequency data, typically used for simulation, is costly to collect and store and therefore is often limited in density and geography. This paper explores the suitability of lower frequency, high availability, connected vehicle data for detailed powertrain simulation. A large data set of 1 Hz trajectories is used to quantify the accuracy loss when simulating energy consumption for conventional, hybrid, and battery electric powertrains using less than 1 Hz data. Techniques to upsample lower frequency drive cycle data in order to increase accuracy are also explored. Median energy consumption errors when simulating energy consumption for a 1/10 Hz trajectory are found to be 3-6% when compared to 1 Hz trajectories. Applying upsampling and interpolation techniques are shown to reduce the simulation errors by roughly 50%. The findings in this work can guide connected vehicle data collection specifications and processing techniques applied when using collected data for powertrain simulation.

ADVANCED PROPULSION SYSTEMS↗

Reliability Assessment of Solid-State Circuit Breakers (CRADA CRD-21-21469, Project 4 Final Report)

Growth in power requirements and the complexity of emerging distribution systems are creating the need for new solid-state products for many applications. The qualification requirements, testing standards and life-cycle management practices of solid-state devices used in protection applications is not established. The project will leverage newly developed accelerated testing capabilities design specifically for protection applications to generate large data sets that will enable advanced digital techniques for data analysis and embedded health monitoring.

14 SOLAR ENERGY↗