Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Development of neural network force fields for corrosion studies

To fully understand the chemistry and physics of corrosion, novel methods of simulation must be developed. One approach is designing machine learning (ML) algorithms integrated with density functional theory to develop adaptive force fields to gain insight into corrosion behavior namely at the surface of metal oxides. Current methods of modeling corrosion are slow due to the computational cost of resolving both reaction mechanics and mass transport processes. Machine learning methods can be implemented to obtain structure-activity relationships at both the molecular and bulk scale while still retaining the accuracy of density functional theory (DFT) and significantly decreasing the time needed for simulations of complex chemical processes in the various environments of corrosion. Multiscale models are needed for corrosion studies to fully understand its processes not only at the atomic length scale (chemical bonding, energies, and forces), but also at the nano and meso length scales (solid-state physics and material science processes). Current methods of study include DFT, molecular dynamics, and Monte Carlo. The limitation of DFT is that only a small number of atoms or molecules can be simulated at that level of theory. Density functional theory is used to study the electronic structure of atoms and molecules, and calculate the force component of each atom. However, these calculations are limited to about 1000 atoms. Custom periodic boundary conditions (PBC) can be used to describe the various environments and defects that affect the atomic forces to produce a large data set from which a training set can be derived. Machine learning can be utilized to overcome the barrier of modeling macroscopic and multi-scale processes from ab initio calculations through the development of adaptive force fields. Local environments determine the atomic forces of a given system, therefore adaptive force fields must be created to produce reliable quantum mechanical calculations. This can be achieved by developing a learning algorithm that uses the mapped atomic forces or fingerprint as an input to produce energies and magnetic moments as output. A systematic approach was used to begin to build a data set in order to accurately describe the atomic forces in various environments. In Figure 4 below, a simple PBC cell of Fe{sub 2}O{sub 3} was first optimized. A surface optimization was performed next, followed by a hydroxylated surface optimization. Once this calculation has converged, the adsorption of halide species to the hydroxylated surface will be investigated. TensorFlow is an open source platform for machine learning developed by Google. Using a high level application program interface (API) such as Keras allows for building and training ML models easily in a number of different environments and languages. For this project, a neural network was developed within Anaconda in Python. Future Work: Further development of reference data set; Refining neural network and learning algorithm; Fingerprinting atomic environment to enable mapping of atomic force components; Choosing appropriate training set from reference data; Learning from training set and enabling non-linear mapping of training set fingerprints and the atomic forces; Estimation of uncertainty to identify ranges of outside applicability; Testing and analysis of molecular dynamic simulations.

36 MATERIALS SCIENCE↗

Opportunities for an Integrated Web-Based Workbench for Data Access and Analysis - 20244

Management of environmental issues can require integration of multiple types of data and information, conducting data analysis and interpretation, and providing data visualization for effective communications. These data elements are important for site management to support regulator interactions and provide defensibility for remedial decisions. Databases and information repositories are core elements of managing data; however, efficient data access and analysis also enable effective site management. The U.S. Department of Energy (DOE) Hanford Site is an example of a complex site with a voluminous quantity of environmental data and a need for efficient site management. Different tiers of data and information tools have been developed and deployed to address site needs. These tools are configured for ready access via the web site interfaces and meet the rigorous quality requirements for environmental site management. Evolving efforts are focused on an integrated platform to meet site environmental management needs. In this platform, users can access site information at multiple levels of detail based on their need and permissions, so that data and associated analyses are presented within the context of the site mission and the user's management or technical needs. This concept is not only applicable at individual sites like Hanford but also applicable at other sites within the DOE complex. An integrated web-based architecture that links data visualization, data analytics, and management tools can provide holistic access to large data sets, minimize complexity, and maximize interactivity and technical communication. (authors)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Nuclear Data Management and Analysis System Plan

The United States Department of Energy Advanced Reactor Technologies Program was formed in Fiscal Year 2015 and encompasses the Next Generation Nuclear Plant Project and Very High Temperature Reactor (VHTR) Program as they were known previously. The VHTR Program was created to support design and licensing of the first VHTR nuclear plant. Data created for and used by the program must be qualified for use, stored in a readily accessible electronic form, categorized to assure the correct data are used, and controlled to prevent data corruption or inadvertent changes. The Nuclear Data Management and Analysis System was designed to support the data needs of the VHTR Program, at the time and now the Advanced Reactor Technologies Program. Since its inception, use of the Nuclear Data Management and Analysis System has expanded to support additional projects and programs with similar requirements for control, analysis, and availability of large data sets.

99 GENERAL AND MISCELLANEOUS↗

Brillouin Sensing with PCA, and PCA-Based Neural Networks for Efficient Temperature Monitoring

This work explores peak estimation techniques in Brillouin Optical Time Domain Analysis (BOTDA), emphasizing both accuracy and efficiency. Euclidean distance measurement method is applied to principal components derived from Brillouin Gain Spectrum data. It offers a major speed advantage being 180 170 times faster than traditional curve fitting methods such as Lorentzian curve fitting, while maintaining similar accuracy. Additionally, a PCA- based neural network model shows significant reduction of peak estimation time compared to Lorentzian fitting. Results show Brillouin frequency shift errors lie under 0.75 MHz in both Euclidean distance-based and neural network-based methods, both of which utilize PCA components. For large data sets and long length fibers, PCA- assisted neural network for peak estimation would be an efficient solution.

Distributed optical fiber sensing↗

Preliminary results from the flight of the Solar Array Module Plasma Interactions Experiment (SAMPIE)

SAMPIE, the Solar Array Module Plasma Interactions Experiment, flew in the Space Shuttle Columbia payload bay as part of the Office of Aeronautics and Space Technology-2 (OAST-2) mission on STS-62, March, 1994. SAMPIE biased samples of solar arrays and space power materials to varying potentials with respect to the surrounding space plasma, and recorded the plasma currents collected and the arcs which occurred, along with a set of plasma diagnostics data. A large set of high quality data was obtained on the behavior of solar arrays and space power materials in the space environment. This paper is the first report on the data SAMPIE telemetered to the ground during the mission. It will be seen that the flight data promise to help determine arcing thresholds, snapover potentials, and floating potentials for arrays and spacecraft in LEO.

Ferguson, Dale C.↗

Preliminary Results from the Flight of the Solar Array Module Plasma Interactions Experiment (SAMPIE)

SAMPIE, the Solar Array Module Plasma Interactions Experiment, flew in the Space Shuttle Columbia payload bay as part of the OAST-2 mission on STS-62, March, 1994. SAMPIE biased samples of solar arrays and space power materials to varying potentials with respect to the surrounding space plasma, and recorded the plasma currents collected and the arcs which occurred, along with a set of plasma diagnostics data. A large set of high quality data was obtained on the behavior of solar arrays and space power materials in the space environment. This paper is the first report on the data SAMPIE telemetered to the ground during the mission. It will be seen that the flight data promise to help determine arcing thresholds, snapover potentials and floating potentials for arrays and spacecraft in LEO.

Ferguson, Dale C.↗

Science Archives in the 21st Century: A NASA LAMBDA Report

Lambda is a thematic data center that focuses on serving the cosmic microwave background (CMB) research community. LAMBDA is an active archive for NASA's Cosmic Background Explorer (COBE) and Wilkinson Microwave Anisotropy Probe (WMAP) mission data sets. In addition, LAMBDA provides analysis software, on-line tools, relevant ancillary data and important web links. LAMBDA also tries to preserve the most important ground-based and suborbital CMB data sets. CMB data is unlike other astrophysical data, consisting of intrinsically diffuse surface brightness photometry with a signal contrast of the order 1 part in 100,000 relative to the uniform background. Because of the extremely faint signal levels, the signal-to-noise ratio is relatively low and detailed instrument-specific knowledge of the data is essential. While the number of data sets being produced is not especially large, those data sets are becoming large and complex. That tendency will increase when the many polarization experiments currently being deployed begin producing data. The LAMBDA experience supports many aspects of the NASA data archive model developed informally over the last ten years-that small focused data centers are often more effective than larger more ambitious collections, for example; that data centers are usually best run by active scientists; that it can be particularly advantageous if those scientists are leaders in the use of the archived data sets; etc. LAMBDA has done some things so well that they might provide lessons for other archives. A lot of effort has been devoted to developing a simple and consistent interface to data sets, for example; and serving all the documentation required via simple 'more' pages and longer explanatory supplements. Many of the problems faced by LAMBDA will also not surprise anyone trying to manage other space science data. These range from persuading mission scientists to provide their data as quickly as possible, to dealing with a high volume of nuisance (spam) messages. Because so many data center problems and solutions are common across individual data centers and disciplines it would be very valuable to establish some new systems of communication - such as informal email lists for administrators and developers. But resources are very limited, so new timeconsuming and inefficient mechanisms - like too-frequent and too-structured meetingsshould be avoided. Although there are great advantages to being small, agile and independent, there are also some areas where science data centers within and without NASA could be better coordinated - for the assignment of persistent identifiers; to encourage the early adoption of useful standards and technologies; etc. Some super-structure to facilitate such coordination might be beneficial as long as it doesn't begin to control the other work of the archives, and become a "methodology police". In this respect the CCSDS "Reference Model for an Open Archive Information System" is a little worrying. It may be that the closer a data center gets to following such a detailed prescription, the less effective it will become. It is much better to have an informal coordination process than a bureaucratic straight-jacket.

Butterworth, P.↗

An interactive environment for the analysis of large Earth observation and model data sets

We propose to develop an interactive environment for the analysis of large Earth science observation and model data sets. We will use a standard scientific data storage format and a large capacity (greater than 20 GB) optical disk system for data management; develop libraries for coordinate transformation and regridding of data sets; modify the NCSA X Image and X Data Slice software for typical Earth observation data sets by including map transformations and missing data handling; develop analysis tools for common mathematical and statistical operations; integrate the components described above into a system for the analysis and comparison of observations and model results; and distribute software and documentation to the scientific community.

Bowman, Kenneth P.↗

An interactive environment for the analysis of large Earth observation and model data sets

We propose to develop an interactive environment for the analysis of large Earth science observation and model data sets. We will use a standard scientific data storage format and a large capacity (greater than 20 GB) optical disk system for data management; develop libraries for coordinate transformation and regridding of data sets; modify the NCSA X Image and X DataSlice software for typical Earth observation data sets by including map transformations and missing data handling; develop analysis tools for common mathematical and statistical operations; integrate the components described above into a system for the analysis and comparison of observations and model results; and distribute software and documentation to the scientific community.

Bowman, Kenneth P.↗

An interactive environment for the analysis of large Earth observation and model data sets

Envision is an interactive environment that provides researchers in the earth sciences convenient ways to manage, browse, and visualize large observed or model data sets. Its main features are support for the netCDF and HDF file formats, an easy to use X/Motif user interface, a client-server configuration, and portability to many UNIX workstations. The Envision package also provides new ways to view and change metadata in a set of data files. It permits a scientist to conveniently and efficiently manage large data sets consisting of many data files. It also provides links to popular visualization tools so that data can be quickly browsed. Envision is a public domain package, freely available to the scientific community. Envision software (binaries and source code) and documentation can be obtained from either of these servers: ftp://vista.atmos.uiuc.edu/pub/envision/ and ftp://csrp.tamu.edu/pub/envision/. Detailed descriptions of Envision capabilities and operations can be found in the User's Guide and Reference Manuals distributed with Envision software.

Bowman, Kenneth P.↗

Compactly‐Supported Nonstationary Kernels for Computing Exact Gaussian Processes on Big Data

The Gaussian process (GP) is a widely used method for analyzing large-scale data sets, including spatio-temporal measurements of nonlinear processes that are now commonplace in the environmental sciences. Traditional implementations of GPs involve stationary kernels (also termed covariance functions) that limit their flexibility, and exact methods for inference that prevent application to data sets with more than about 10,000 points. Modern approaches to address stationarity assumptions generally fail to accommodate large data sets, while all attempts to address scalability focus on approximating the Gaussian likelihood, which can involve subjectivity and lead to inaccuracies. In this work, we explicitly derive an alternative kernel that can discover and encode both sparsity and nonstationarity. We embed the kernel within a fully Bayesian GP model and leverage high-performance computing resources to enable the analysis of massive data sets. We demonstrate the favorable performance of our novel kernel relative to existing exact and approximate GP methods across a variety of synthetic data examples. Furthermore, we conduct space–time prediction based on more than 1 million measurements of daily maximum temperature and verify that our results outperform state-of-the-art methods in the Earth sciences. More broadly, having access to exact GPs that use ultra-scalable, sparsity-discovering, nonstationary kernels allows GP methods to truly compete with a wide variety of machine learning methods.

Gaussian processes↗

The Use of a Satellite Climatological Data Set to Infer Large Scale Three Dimensional Flow Characteristics

Ever since the first satellite image loops from the 6.3 micron water vapor channel on the METEOSAT-1 in 1978, there have been numerous efforts (many to a great degree of success) to relate the water vapor radiance patterns to familiar atmospheric dynamic quantities. The realization of these efforts is becoming evident with the merging of satellite derived winds into predictive models (Velden et al., 1997; Swadley and Goerss, 1989). Another parameter that has been quantified from satellite water vapor channel measurements is upper tropospheric relative humidity (UTH) (e.g., Soden and Bretherton, 1996; Schmetz and Turpeinen, 1988). These humidity measurements, in turn, can be used to quantify upper tropospheric water vapor and its transport to more accurately diagnose climate changes (Lerner et al., 1998; Schmetz et al. 1995a) and quantify radiative processes in the upper troposphere. Also apparent in water vapor imagery animations are regions of subsiding and ascending air flow. Indeed, a component of the translated motions we observe are due to vertical velocities. The few attempts at exploiting this information have been met with a fair degree of success. Picon and Desbois (1990) statistically related Meteosat monthly mean water vapor radiances to six standard pressure levels of the European Centre for Medium Range Weather Forecast (ECMWF) model vertical velocities and found correlation coefficients of about 0.50 or less. This paper presents some preliminary results of viewing climatological satellite water vapor data in a different fashion. Specifically, we attempt to infer the three dimensional flow characteristics of the mid- to upper troposphere as portrayed by GOES VAS during the warm ENSO event (1987) and a subsequent cold period in 1998.

Lerner, Jeffrey A.↗

A Comparison of Ffowcs Williams-Hawkings Solvers for Airframe Noise Applications

This paper presents a comparison between two implementations of the Ffowcs Williams and Hawkings equation for airframe noise applications. Airframe systems are generally moving at constant speed and not rotating, so these conditions are used in the current investigation. Efficient and easily implemented forms of the equations applicable to subsonic, rectilinear motion of all acoustic sources are used. The assumptions allow the derivation of a simple form of the equations in the frequency-domain, and the time-domain method uses the restrictions on the motion to reduce the work required to find the emission time. The comparison between the frequency domain method and the retarded time formulation reveals some of the advantages of the different approaches. Both methods are still capable of predicting the far-field noise from nonlinear near-field flow quantities. Because of the large input data sets and potentially large numbers of observer positions of interest in three-dimensional problems, both codes utilize the message passing interface to divide the problem among different processors. Example problems are used to demonstrate the usefulness and efficiency of the two schemes.

Lockard, David P.↗

Prediction and Experimental Verification of Electrolyte Solvation Structure from an OMol25-Trained Interatomic Potential

A molecular-level understanding of electrolyte solvation structure and ion–ion correlations is critical to developing next-generation battery chemistries. Atomistic simulation capabilities with sufficient accuracy, speed, and transferability to deliver reliable structural insights while avoiding arduous system-specific reparameterization are thus highly desirable. Machine learning interatomic potentials (MLIPs) trained on large, chemically diverse data sets are revolutionizing computational chemistry, enabling molecular dynamics simulations of battery electrolytes with near-DFT accuracy over 10,000× faster than DFT. While previous MLIP training data sets with suitable elemental coverage for electrolytes have been based on inorganic materials, the Open Molecules 2025 (OMol25) data set provides large-scale molecular DFT MLIP training data with broad elemental coverage and specifically samples tens of millions of electrolyte configurations. Here, we integrate computational modeling with experimental validation to systematically assess the ability of large-scale MLIPs pretrained on materials data or on OMol25 to accurately resolve nanoscale structural organization and ion-solvation characteristics in Na-ion battery electrolytes across diverse physicochemical conditions and compositional regimes. We find that the OMol25-trained Universal Model of Atoms (UMA-OMol) predicts experimentally measured densities and X-ray structure factors in substantially better agreement compared to state-of-the-art models trained only on inorganic materials data. Using UMA-OMol, we further analyze systematic trends in solvation structure as a function of cation identity, anion chemistry, salt concentration, and solvent topology. We observe that increasing system temperature amplifies the heterogeneity within the solvation environment, perturbing cation–solvent interactions and promoting the formation of contact ion pairs (CIPs). Moreover, subtle variations in the solvent topology of glyme-based electrolytes cause pronounced changes in ion correlations and solvation structure. The experimental agreement and microscopic insights shown here position OMol25-trained MLIPs as a practical route to predictive, high-throughput electrolyte simulations beyond the limits of classical force fields and direct DFT molecular dynamics, serving as a powerful tool for accelerating the design of next-generation Na-ion battery electrolytes and beyond.

MLIPs↗