Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Redox transistors based on TiO 2 for analogue neuromorphic computing

The ability to train deep neural networks on large data sets have made significant impacts onto artificial intelligence, but consume significant amounts of energy due to the need to move information from memory to logic units. In-memory "neuromorphic" computing presents an alternative framework that processes information directly on memory elements. In-memory computing has been limited by the poor performance of the analogue information storage element, often phase-change memory or memristors. To solve this problem, we developed two types of "redox transistors" using TiO 2 (anatase) which stores analogue information states through the electrochemical concentration of dopants in the crystal. The first type of redox transistor uses lithium as the electrochemical dopant ion, and its key advantage is low operating voltage. The second uses oxygen vacancies as the dopant, which is CMOS compatible and can retain state even when scaled to nanosized dimensions. Both devices offer significant advantages in terms of predictable analogue switching over conventional filamentary-based devices, and provide a significant advance in developing materials and devices for neuromorphic computing.

36 MATERIALS SCIENCE↗

Identifying Suspect Instrument Intervals Using Midnight Noise Time Histories

With an intense work ethic, and high levels of devotion to their craft, seismologists continue to battle the elements and wrestle with complex logistics in their efforts to field instrumentation over the entire globe. We are fortunate to see data volumes accelerating in size, and quality continuing to improve; for example, universal timing issues can now be considered rare. However, instrument response issues are still common, and can hinder research that seeks to understand amplitudes of seismic wavefields. We are interested in using the accumulations of seismic data to develop models that will predict high frequency (0.2-20 Hz, and higher) signal amplitudes for explosion monitoring purposes over broad areas, which requires large data sets and extensive quality control. To identify response and station health issues, we have collected noise time histories for global seismic data, focusing on measurements near (but not restricted to) midnight to eliminate diurnal variations, and have manually determined time intervals that appear inconsistent with background behavior. We assign descriptive labels, but do not attempt to diagnose causes. We use these intervals to discard data. To date, we have examined 39,260 channels from 11,105 stations, heavily weighted toward IRIS holdings, dates through 2017 (depending on station, roughly the date we started our manual review), bands between 1 and 8 Hz, finding 24,733 anomalous time intervals. The great majority (90%) of these intervals appear to be shifts of constant offset, often bounded by times of known instrument changes, likely the result of poor documentation of response parameters at one of many stages between the field and the plotter. We hope these results can be of use to our colleagues, and would encourage community efforts to diagnose anomalous behavior, and fix poor responses. We also hope these results will support automation efforts, including application of supervised learning techniques.

47 OTHER INSTRUMENTATION↗

Use of the Tool to Support Renewable Energy Auctions Processes (RE Data Explorer)

Renewable energy auctions are now a common competitive approach to procure low-cost renewable power around the world. Ensuring a successful auction process increasingly depends on the capabilities of auction designers and participants to identify actionable and defensible insights from large data sets (on renewable energy resources and complementary data) to both attract potential investors and address stakeholder concerns. The Renewable Energy (RE) Data Explorer is a user-friendly geospatial analysis tool for analyzing renewable energy potential and informing decisions. Developed by the National Renewable Energy Laboratory (NREL) and supported by the U.S. Agency for International Development (USAID), RE Data Explorer performs visualization and analysis of renewable energy potential that can be customized for different scenarios. RE Data Explorer can support prospecting, integrated planning, policymaking, and other decision-making activities to accelerate renewable energy deployment. The broader RE Explorer website provides guidance and information to link the RE Data Explorer geospatial analysis tool to key decision areas. This document provides information on how the RE Data Explorer can be used to support renewable energy auction processes.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Rheo-Structural Spectroscopy: Fingerprinting the In Situ Response of Fluids to Arbitrary Flow Fields

The objectives of this project were to develop new sample environments, measurement methodologies and associated modeling tools for characterizing the structural response to arbitrarily complex processing flows using small angle scattering, and to apply these new tools for understanding the fundamental physics governing the structuring of anisotropic particulate and polymeric materials under flow histories and conditions relevant to industrial processing flows. The research resulted in the development and implementation of a new sample environment, the fluidic four roll mill (FFoRM), for in situ small angle neutron and X-ray scattering (SANS/SAXS) measurements. These measurements are capable of generating large data sets that “fingerprint” how a complex fluid responds to a wide range of flow histories involving time variations in deformation type and rate. New modeling tools were developed to extract detailed microstructural information from such data sets, including orientation distribution functions and interparticle correlation functions, as well as reduced-order parametric descriptors of these high-dimensional functions that can be used to readily map, visualize and interpret a fluid’s structural response to its flow history. These new tools were applied to a range of model materials involving elongated particle suspensions in order to provide new insights into the physics of how flow couples with orientational and structural order in complex flows, particularly under non-dilute conditions for which no accurate theories currently exist. Using these investigations, we elucidated a number of new insights into the fundamental phenomena driving such process-structure-property relationships. These findings provide guidance for the further development of rheological models, and ultimately can inform the rational and model-based design of flow processes to achieve optimized orientational ordering that is key to the properties and function of a wide range of energy-relevant materials.

36 MATERIALS SCIENCE↗

Simultaneous mitigation of density and energy errors in approximate DFT for transition metal chemistry (Final Technical Report)

There were three major goals and objectives of this project: 1) Develop tools to understand density-driven errors in transition metal complexes, 2) Evaluate and minimize energy delocalization error and static correlation error through judicious functional choice, and 3) applying this workflow to machine-learning accelerated screening of redox couples. Over the reporting period, we developed a framework for eliminating flat plane errors. We introduced fully non-empirical coefficients. We demonstrated the approach on both molecules and solids. We investigated and eliminated density driven errors and demonstrated their impact on potential energy surfaces. We trained machine learning models both in a method-dependent fashion and to predict errors in method accuracy. We built large data sets of small molecule energetics and multi-reference character.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Data Science and Machine Learning in Education

The growing role of data science (DS) and machine learning (ML) in high-energy physics (HEP) is well established and pertinent given the complex detectors, large data, sets and sophisticated analyses at the heart of HEP research. Moreover, exploiting symmetries inherent in physics data have inspired physics-informed ML as a vibrant sub-field of computer science research. HEP researchers benefit greatly from materials widely available materials for use in education, training and workforce development. They are also contributing to these materials and providing software to DS/ML-related fields. Increasingly, physics departments are offering courses at the intersection of DS, ML and physics, often using curricula developed by HEP researchers and involving open software and data used in HEP. In this white paper, we explore synergies between HEP research and DS/ML education, discuss opportunities and challenges at this intersection, and propose community activities that will be mutually beneficial.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

IDAES Enterprise: Generation Expansion Planning with Enhanced Requirements for Capacity Adequacy Under Renewable Intermittency

Achieving net zero carbon emissions likely requires future power systems to integrate new, flexible energy technologies to accommodate higher levels of capacity from variable renewable energy sources. To determine the optimal deployment of new electricity capacity and to study the likelihood of deployments of new energy technologies, an expansion planning model has been developed as part of the IDAES-Enterprise suite of grid models. The Generation Expansion Planning (GEP) model is a multi-period model in which investment decisions occur yearly, and a Unit Commitment (UC) problem is examined on an hourly timescale. To reduce computational complexity of the GEP model, the UC problem is solved for average “representative days” which leaves out extreme, but relatively common, scenarios in which low renewable generation occurs, leaving the system with inadequacy in capacity. The IDAES-Enterprise GEP model has been modified to include these extreme scenarios while keeping the model reasonably tractable. Specifically, a lazy constraint technique was implemented to check for capacity adequacy on an hourly basis over a large data set of aligned load-wind-solar profiles. As a vast majority of the capacity constraints will not be violated, the technique lowers computational expense by searching for violated capacity constraints over an “iterative manner,” adding those infeasible constraints back into the model. Results on a test case of the Southwest Power Pool shows that the lazy constraint technique significantly reduces retirements and increases installments of natural gas combined cycles and flexible natural gas units. It also reduces some retirements of coal units. These modifications provide a more reasonable estimation of required dispatchable power generation capacity to ensure feasibility during peak net load.

Liu, Peng↗

Measurement of cross sections for mesonless charged-current muon neutrino interactions on argon with and without protons at MicroBooNE

urrent and upcoming neutrino experiments at Fermilab will rely on liquid argon time projection chambers (LArTPCs) as the primary detector technology. To reach the ambitious precision required for their physics goals, a detailed understanding of neutrino-argon scattering must be achieved. For the present Short-Baseline Neutrino program, the dominant reaction channel is charged-current muon neutrino interactions leading to mesonless final states with and without protons. Using a large data set of neutrino interactions recorded by the MicroBooNE LArTPC detector at Fermilab, we report progress towards new measurements of differential cross sections in this channel. The analysis leverages multiple event reconstruction paradigms and observables specifically chosen to maximize sensitivity to nuclear effects of interest for improving neutrino scattering models. Related previous cross-section results from MicroBooNE and a preview of what to expect from the new analysis will both be featured in this presentation.

Liu, Liang [Fermilab]↗

Gamma Ray Source Localization for Time Projection Chamber Telescopes Using Convolutional Neural Networks

Diverse phenomena such as positron annihilation in the Milky Way, merging binary neutron stars, and dark matter can be better understood by studying their gamma ray emission. Despite their importance, MeV gamma rays have been poorly explored at sensitivities that would allow for deeper insight into the nature of the gamma emitting objects. In response, a liquid argon time projection chamber (TPC) gamma ray instrument concept called GammaTPC has been proposed and promises exploration of the entire sky with a large field of view, large effective area, and high polarization sensitivity. Optimizing the pointing capability of this instrument is crucial and can be accomplished by leveraging convolutional neural networks to reconstruct electron recoil paths from Compton scattering events within the detector. In this investigation, we develop a machine learning model architecture to accommodate a large data set of high fidelity simulated electron tracks and reconstruct paths. We create two model architectures: one to predict the electron recoil track origin and one for the initial scattering direction. We find that these models predict the true origin and direction with extremely high accuracy, thereby optimizing the observatory’s estimates of the sky location of gamma ray sources.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A Comprehensive Characterization of Galaxy-cool CGM Connections at z < 0.4 with DESI Year 1 Data

We investigate the relationships between the cool circumgalactic medium (CGM), traced by Ca II absorption lines, and galaxy properties at z < 0.4 using ∼900,000 galaxy–quasar pairs within 200 kpc from the Year 1 data of the Dark Energy Spectroscopic Instrument (DESI). This large data set enables us to obtain composite spectra with sensitivity reaching to the mÅ level and to explore the Ca II absorption as a function of stellar mass, star formation rate (SFR), redshift, and galaxy types, including active galactic nuclei (AGNs). Our results show a positive correlation between the absorption strength and stellar mass of star-forming galaxies with $\langle$$W$$^{Ca II}_{0}$$\rangle$ α $M$$^{0.5}_{*}$ over 3 orders of magnitude in stellar mass from ∼10 8 to 10 11 M ⊙ , while such a mass dependence is weaker for quiescent galaxies. At a fixed mass, Ca II absorption is stronger around star-forming galaxies than quiescent ones especially within impact parameters <30 kpc. Among star-forming galaxies, the Ca II absorption further correlates with SFR, following ∝SFR 0.3 . However, in contrast to the results at higher redshifts, stronger absorption is not preferentially observed along the minor axis of star-forming galaxies, indicating a possible redshift evolution of CGM dynamics resulting from galactic feedback. Moreover, no significant difference between the properties of the cool gas around AGNs and galaxies is detected. Finally, we measure the absorption profiles with respect to the virial radius of dark matter halos and show that the total Ca II mass in the CGM is comparable to the Ca mass in the ISM of galaxies.

Circumgalactic medium↗

Vehicle Powertrain Simulation Accuracy for Various Drive Cycle Frequencies and Upsampling Techniques

As connected and automated vehicle technologies emerge and proliferate, lower frequency vehicle trajectory data is becoming more widely available. In some cases, entire fleets are streaming position, speed, and telemetry at sample rates of less than 10 seconds. This presents opportunities to apply powertrain simulators such as the National Renewable Energy Laboratory's Future Automotive Systems Technology Simulator to model how advanced powertrain technologies would perform in the real world. However, connected vehicle data tends to be available at lower temporal frequencies than the 1-10 Hz trajectories that have typically been used for powertrain simulation. Higher frequency data, typically used for simulation, is costly to collect and store and therefore is often limited in density and geography. This paper explores the suitability of lower frequency, high availability, connected vehicle data for detailed powertrain simulation. A large data set of 1 Hz trajectories is used to quantify the accuracy loss when simulating energy consumption for conventional, hybrid, and battery electric powertrains using less than 1 Hz data. Techniques to upsample lower frequency drive cycle data in order to increase accuracy are also explored. Median energy consumption errors when simulating energy consumption for a 1/10 Hz trajectory are found to be 3-6% when compared to 1 Hz trajectories. Applying upsampling and interpolation techniques are shown to reduce the simulation errors by roughly 50%. The findings in this work can guide connected vehicle data collection specifications and processing techniques applied when using collected data for powertrain simulation.

ADVANCED PROPULSION SYSTEMS↗

Reliability Assessment of Solid-State Circuit Breakers (CRADA CRD-21-21469, Project 4 Final Report)

Growth in power requirements and the complexity of emerging distribution systems are creating the need for new solid-state products for many applications. The qualification requirements, testing standards and life-cycle management practices of solid-state devices used in protection applications is not established. The project will leverage newly developed accelerated testing capabilities design specifically for protection applications to generate large data sets that will enable advanced digital techniques for data analysis and embedded health monitoring.

14 SOLAR ENERGY↗

Improved Access and Analysis of Data Provides New Opportunities for Smart Cities

Making smart, informed, data-driven energy decisions for cities requires large amounts of high-value data and analysis. Cities’ energy data has traditionally been difficult to access and costly to store and manage for city administrators. When data is available, many cities lack the expertise to properly analyze and model large data sets. NREL’s transparency and knowledge about data availability, integration, and interpretation is critical to decision-making for cities.

cities' data↗

PV Reliability Lessons from 100,000 Systems

Despite the importance of reliability to the cost competitiveness of PV, large data sets enabling high-level investigation of the technology’s performance in the field are relatively scarce. Dirk C. Jordan, Chris Deline, Bill Marion and Teresa Barnes of the National Renewable Energy Laboratory, and Mark Bolinger of the Lawrence Berkeley National Laboratory study a unique data set of 100,000 PV systems in the US, drawing out tips for better reliability that have relevance to other parts of the world.

41 EE - Solar Energy Technologies Office (EE-4S)↗

On the Use of Smart Meter Data to Estimate the Voltage Magnitude on the Primary Side of Distribution Service Transformers: Preprint

This paper develops a novel method to estimate the voltage magnitude on the primary side of distribution service transformers. The proposed method relies exclusively on smart meters, and therefore it is fully data-driven. This is an important feature because electric utilities have detailed models of only the primary network--that is, the network between the distribution substation and the primary side of service transformers that are installed closer to end-customer sites. The network that connects the secondary side of service transformers to end-customer sites, referred to as the secondary network, is simply represented by a lumped load. For each secondary network, the proposed method uses data acquired from only 2 smart meters: the closest and the farthest--in the sense of electrical distance--from the service transformer. As a reference to this feature, the proposed method is named SM2Vp. To our knowledge, this is the first time a method is shown to provide actionable information for real-time operation and control of power distribution grids using only two smart meters per secondary network. This is important because utilities have experienced barriers in managing and using large data sets for real-time operation and control. SM2Vp is primarily intended to provide pseudo-measurements for distribution system state estimation, but it can also be used directly for voltage control schemes. The performance of SM2Vp is demonstrated by numerical simulations carried out on three secondary network synthetic models and by using field data provided by a utility partner serving customers in southwestern California. A maximum relative error of approximately 3.9% or less is observed for the primary voltage magnitude estimates in all numerical experiments.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Advancing the Theory of Nuclear Data Evaluations [Abstract]

We present recent advances in the R-matrix formalism as well as the Bayesian evaluation framework for improved nuclear data evaluations. The advances in the R matrix formalism include: 1) direct processes, 2) doorway, as well as multistep, processes, and 3) various forms of the Reich-Moore approximation for eliminated capture channels. Furthermore, to address unreasonably small posterior uncertainties often encountered in nuclear data evaluations of large data sets using the conventional form of the Bayes’ theorem, we introduce imperfections (of the data or the model) as a formal evaluation tool for taming the evaluated uncertainties in harmony with Bayes’ theorem. These theoretical advances were motivated by the nuclear data evaluations of differential resolved resonance cross section data using the code SAMMY, as well as the integral benchmark experiments using the SCALE code system, being performed at Oak Ridge National Laboratory for the Nuclear Criticality Safety Program. Some pedagogical applications of the new formalism, as well as a snapshot of the SAMMY modernization efforts, will be presented.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Development of neural network force fields for corrosion studies

To fully understand the chemistry and physics of corrosion, novel methods of simulation must be developed. One approach is designing machine learning (ML) algorithms integrated with density functional theory to develop adaptive force fields to gain insight into corrosion behavior namely at the surface of metal oxides. Current methods of modeling corrosion are slow due to the computational cost of resolving both reaction mechanics and mass transport processes. Machine learning methods can be implemented to obtain structure-activity relationships at both the molecular and bulk scale while still retaining the accuracy of density functional theory (DFT) and significantly decreasing the time needed for simulations of complex chemical processes in the various environments of corrosion. Multiscale models are needed for corrosion studies to fully understand its processes not only at the atomic length scale (chemical bonding, energies, and forces), but also at the nano and meso length scales (solid-state physics and material science processes). Current methods of study include DFT, molecular dynamics, and Monte Carlo. The limitation of DFT is that only a small number of atoms or molecules can be simulated at that level of theory. Density functional theory is used to study the electronic structure of atoms and molecules, and calculate the force component of each atom. However, these calculations are limited to about 1000 atoms. Custom periodic boundary conditions (PBC) can be used to describe the various environments and defects that affect the atomic forces to produce a large data set from which a training set can be derived. Machine learning can be utilized to overcome the barrier of modeling macroscopic and multi-scale processes from ab initio calculations through the development of adaptive force fields. Local environments determine the atomic forces of a given system, therefore adaptive force fields must be created to produce reliable quantum mechanical calculations. This can be achieved by developing a learning algorithm that uses the mapped atomic forces or fingerprint as an input to produce energies and magnetic moments as output. A systematic approach was used to begin to build a data set in order to accurately describe the atomic forces in various environments. In Figure 4 below, a simple PBC cell of Fe{sub 2}O{sub 3} was first optimized. A surface optimization was performed next, followed by a hydroxylated surface optimization. Once this calculation has converged, the adsorption of halide species to the hydroxylated surface will be investigated. TensorFlow is an open source platform for machine learning developed by Google. Using a high level application program interface (API) such as Keras allows for building and training ML models easily in a number of different environments and languages. For this project, a neural network was developed within Anaconda in Python. Future Work: Further development of reference data set; Refining neural network and learning algorithm; Fingerprinting atomic environment to enable mapping of atomic force components; Choosing appropriate training set from reference data; Learning from training set and enabling non-linear mapping of training set fingerprints and the atomic forces; Estimation of uncertainty to identify ranges of outside applicability; Testing and analysis of molecular dynamic simulations.

36 MATERIALS SCIENCE↗

Opportunities for an Integrated Web-Based Workbench for Data Access and Analysis - 20244

Management of environmental issues can require integration of multiple types of data and information, conducting data analysis and interpretation, and providing data visualization for effective communications. These data elements are important for site management to support regulator interactions and provide defensibility for remedial decisions. Databases and information repositories are core elements of managing data; however, efficient data access and analysis also enable effective site management. The U.S. Department of Energy (DOE) Hanford Site is an example of a complex site with a voluminous quantity of environmental data and a need for efficient site management. Different tiers of data and information tools have been developed and deployed to address site needs. These tools are configured for ready access via the web site interfaces and meet the rigorous quality requirements for environmental site management. Evolving efforts are focused on an integrated platform to meet site environmental management needs. In this platform, users can access site information at multiple levels of detail based on their need and permissions, so that data and associated analyses are presented within the context of the site mission and the user's management or technical needs. This concept is not only applicable at individual sites like Hanford but also applicable at other sites within the DOE complex. An integrated web-based architecture that links data visualization, data analytics, and management tools can provide holistic access to large data sets, minimize complexity, and maximize interactivity and technical communication. (authors)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗