Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Modern data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Preventing Reverse Engineering of Critical Industrial Data with DIOD

Business analytics augmented by artificial intelligence and machine learning (AI/ML) have revolutionized the role of data in the modern world. In recent years, businesses have incorporated data into their decision-making process for better prediction, risk-assessment, content creation, etc. While such businesses often seek to leverage the full use of their data through third-party AI/ML services, they are often hampered by the risks of data leaks, reverse-engineering, stolen technology, etc. that often have disastrous consequences for businesses and their stakeholders alike. Thus, there arises a need for data masking prior to its transmission that obfuscates proprietary information while preserving the information relevant for AI/ML applications. In order to meet the needs of industrial data which are significantly different from those of data warehouses, previous work proposed an efficient time and space-scalable data masking paradigm known as the deceptive infusion of data (DIOD) methodology. The present work expands upon this work by leveraging existing reverse-engineering capabilities to facilitate the decomposition of industrial data into its proprietary and AI/ML-relevant parts, referred to as fundamental and inference metadata respectively. Both sets of metadata are further obfuscated in accordance with the DIOD methodology to create the DIOD rendition of the industrial data, which is rendered immune to reverse-engineering by discarding proprietary information and only preserving AI/ML-relevant information. Additionally, constraints of the original DIOD manuscript are relaxed using mutual information by configuring the methodology to the target AI/ML application to unlock the full potential of the DIOD methodology. As an example, data from a nuclear reactor is transformed into that from a nonlinear spring-mass system with different levels of data masking as required by the generic system and the target application.

97 MATHEMATICS AND COMPUTING↗

Hadoop for High-Performance Climate Analytics: Use Cases and Lessons Learned

Scientific data services are a critical aspect of the NASA Center for Climate Simulations mission (NCCS). Hadoop, via MapReduce, provides an approach to high-performance analytics that is proving to be useful to data intensive problems in climate research. It offers an analysis paradigm that uses clusters of computers and combines distributed storage of large data sets with parallel computation. The NCCS is particularly interested in the potential of Hadoop to speed up basic operations common to a wide range of analyses. In order to evaluate this potential, we prototyped a series of canonical MapReduce operations over a test suite of observational and climate simulation datasets. The initial focus was on averaging operations over arbitrary spatial and temporal extents within Modern Era Retrospective- Analysis for Research and Applications (MERRA) data. After preliminary results suggested that this approach improves efficiencies within data intensive analytic workflows, we invested in building a cyber infrastructure resource for developing a new generation of climate data analysis capabilities using Hadoop. This resource is focused on reducing the time spent in the preparation of reanalysis data used in data-model inter-comparison, a long sought goal of the climate community. This paper summarizes the related use cases and lessons learned.

analytics↗

Selection of Global Climate Model Data for Downscaling With Generative Machine Learning and Use in the Power Planning for Alignment of Climate and Energy Systems Project

The range of results from climate models and scenarios is important to the understanding of uncertainty in power planning analysis. A U.S. Department of Energy-funded analytic project called Power Planning for Alignment of Climate and Energy Systems is developing data and analytic methods to reflect the effects of climate change on key variables for power system planning, as part of the Grid Modernization Lab Consortium. This project will select and prepare global climate model results for use in power system planning models. A related report (Evaluation of Global Climate Models for Use in Energy Analysis) assesses the performance of various global climate models from the Coupled Model Intercomparison Project Phase 6 data archive for their historical skill with respect to energy system performance and for their future projections under multiple climate change scenarios. Building from that report, we describe the selection of a climate scenario (Shared Socioeconomic Pathway [SSP] 2-4.5) and five climate models: TaiESM1, EC-Earth3-CC, GFDL-CM4, EC-Earth3-Veg, and MPI-ESM1-2-HR. We describe the model selection criteria, which were based on the quality of the match between model results under historical conditions and on the representation of the range of future values for several variables. These results will be downscaled via an open-source generative machine learning method called Super-Resolution for Renewable Energy Resource Data with Climate Change Impacts.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Materials Engineering of Violin Soundboards by Stradivari and Guarneri

Abstract We investigated the material properties of Cremonese soundboards using a wide range of spectroscopic, microscopic, and chemical techniques. We found similar types of spruce in Cremonese soundboards as in modern instruments, but Cremonese spruces exhibit unnatural elemental compositions and oxidation patterns that suggest artificial manipulation. Combining analytical data and historical information, we may deduce the minerals being added and their potential functions—borax and metal sulfates for fungal suppression, table salt for moisture control, alum for molecular crosslinking, and potash or quicklime for alkaline treatment. The overall purpose may have been wood preservation or acoustic tuning. Hemicellulose fragmentation and altered cellulose nanostructures are observed in heavily treated Stradivari specimens, which show diminished second‐harmonic generation signals. Guarneri's practice of crosslinking wood fibers via aluminum coordination may also affect mechanical and acoustic properties. Our data suggest that old masters undertook materials engineering experiments to produce soundboards with unique properties.

Su, Cheng‐Kuan↗

Materials Engineering of Violin Soundboards by Stradivari and Guarneri

Abstract We investigated the material properties of Cremonese soundboards using a wide range of spectroscopic, microscopic, and chemical techniques. We found similar types of spruce in Cremonese soundboards as in modern instruments, but Cremonese spruces exhibit unnatural elemental compositions and oxidation patterns that suggest artificial manipulation. Combining analytical data and historical information, we may deduce the minerals being added and their potential functions—borax and metal sulfates for fungal suppression, table salt for moisture control, alum for molecular crosslinking, and potash or quicklime for alkaline treatment. The overall purpose may have been wood preservation or acoustic tuning. Hemicellulose fragmentation and altered cellulose nanostructures are observed in heavily treated Stradivari specimens, which show diminished second‐harmonic generation signals. Guarneri's practice of crosslinking wood fibers via aluminum coordination may also affect mechanical and acoustic properties. Our data suggest that old masters undertook materials engineering experiments to produce soundboards with unique properties.

36 MATERIALS SCIENCE↗

A study of the effects of a cubic nonlinearity on a modern modal identification technique

The effect of a geometric nonlinearity on the Ibrahim Time Domain (ITD) modal data analysis technique has been studied using two analytically derived models and one laboratory model. Response data for the three models were analyzed by the ITD method. Indicators of nonlinear response were found which include harmonically related frequencies with repetitive mode shapes, clusters of frequencies in a narrow band around each harmonic, and variations in frequency with amplitude of oscillation. Also,for the cases studied, the presence of a nonlinearity has no detrimental effect on identifying linear responses. A potential for applying the algorithm to the identification of a certain class of nonlinear system was indicated.

Horta, L. G.↗

The University of Arizona program in solid propellants

The University of Arizona program is aimed at introducing scientific rigor to the predictability and quality assurance of composite solid propellants. Two separate approaches are followed: to use the modern analytical techniques to experimentally study carefully controlled propellant batches to discern trends in mixing, casting, and cure; and to examine a vast bank of data, that has fairly detailed information on the ingredients, processing, and rocket firing results. The experimental and analytical work is described briefly. The principle findings were that: (1) pre- (dry) blending of the coarse and fine ammonium perchlorate can significantly improve the uniformity of mixing; (2) the Fourier transformed IR spectra of the uncured and cured polymer have valuable data on the state of the fuel; (3) there are considerable non-uniformities in the propellant slurry composition near the solid surfaces (blades, walls) compared to the bulk slurry; and (4) in situ measurements of slurry viscosity continuously during mixing can give a good indication of the state of the slurry. Several important observations in the study of the data bank are discussed.

Ramohalli, Kumar↗

Dare Mighty Things

Winston Churchill once said, “To improve is to change; to be perfect is to change often.” JPL’s Property Accountability objective is to provide superior services related to property accountability, reutilization, and disposition. JPL must report property reutilized and disposed through either sales, donation, or scrap. Implementing the JPL designed and built Property Information Reporting System (PIRS), opened a new perspective with our data and what are we reporting to NASA. With this new depth of visibility, we self-implemented as assessment to validate the fiduciary and stewardship responsibilities of what is being reporting to NASA. JPL strives to perform at a level beyond the basic primary expectations of our NASA requirements. JPL developed the PIRS - which rolled out in 2018 - to simplify the delivery of JPL’s Personal Property and Equipment (PP&E) reporting: PIRS takes the raw data from Oracle to generate the NASA Form (NF) 1018 and Contractor Held Asset Reporting System (CHATS) Reports, and to meet the requirements established for AS9100 compliance. The data presented from PIRS is exportable and used to analyze our property records. JPL is also embarking on an effort to define the JPL of the future with “Enterprise 2.0”. Enterprise 2.0 is the driving force of JPL’s strategy to digitally transform the Laboratory from the current legacy processes and systems to a modern, integrated, information-driven highway of business transactions. This should include automation of routine processes and data-related tasks, integration of data and systems, advanced search and analytics, and improved information sharing and collaboration. As JPL’s evolving landscape of digital technologies advance to meet the future, JPL Property Accountability is making strides to surpass these expectations. For every cause there is an effect: a revealing of something that aids in the further development and transformation of JPL best business practices.

Sucher, Jay M↗

The Land Surface Data Toolkit (LDT v7.2) - A Data Fusion Environment for Land Data Assimilation Systems

The effective applications of land surface models (LSMs) and hydrologic models pose a varied set of data input and processing needs, ranging from ensuring consistency checks to more derived data processing and analytics. This article describes the development of the Land surface Data Toolkit (LDT), which is an integrated framework designed specifically for processing input data to execute LSMs and hydrological models. LDT not only serves as a preprocessor to the NASA Land Information System (LIS), which is an integrated framework designed for multi-model LSM simulations and data assimilation (DA) integrations, but also as a land-surface-based observation and DA input processor. It offers a variety of user options and inputs to processing datasets for use within LIS and stand-alone models. The LDT design facilitates the use of common data formats and conventions. LDT is also capable of processing LSM initial conditions and meteorological boundary conditions and ensuring data quality for inputs to LSMs and DA routines. The machine learning layer in LDT facilitates the use of modern data science algorithms for developing data-driven predictive models. Through the use of an object-oriented framework design, LDT provides extensible features for the continued development of support for different types of observational datasets and data analytics algorithms to aid land surface modeling and data assimilation.

droughts and floods↗

Genesis Solar Wind Sample Curation Documentation

Introduction: A scientist with experience as a sample science analyst, provider of flight hardware for multiple missions, and senior engineer in an ISO 2000-rated manufacturing plant has described the timeline of key participants in any PI-led sample return mission, the breadth of the organizations involved [1,2], and, of interest to this meeting, choosing the types of data to preserve and issues of future data accessibility. This work broadens that perspective by giving similar lessons from Genesis sample curation point-of-view. Curation participation regarding data gathering was part of the mission review process from the beginning. Genesis’ story illustrates outcome of several choices about types of data to record and preserve. Precision analysis of solar wind atoms captured in pure, ultraclean substrates is the driving science goal; therefore, detailed documentation was captured from all mission and curation phases and from investigator laboratories because these processes affect the final analytical results [3]. Pre-flight: Design and fabrication of the spacecraft. Like many modern small sample return missions, Genesis was a tightly managed team integrated across science, engineering and curation. Communication across the team was excellent, and, for the most part, the hands-on engineering technicians understood the impacts of “small choices” they routinely make, and the eyes-on oversight of manufacturing processes by scientists was mindful of details. The payload was designed by the Jet Propulsion Laboratory and the spacecraft by Lockheed Martin. Solar wind collectors and instruments were fabricated by multiple vendors and laboratories. The main portion of the payload was assembled at JSC. Fabrication procedures and contamination-control data (with witness coupons) were stored primarily at JSC. The original composition, dimensions and configuration of components, results of thermal testing, etc. are still needed for interpretation of analytical data. At times, these must be estimated from secondary information acquired pre-flight. Moreover, some files (e.g., original 3-D models and early Powerpoint) cannot be opened using software. Archived curation data includes 2-D drawings, material usage lists, QA documentation and analyses of consumables used during fabrication. Important chemical information still resides in archived hardware, paints and lubricants, material coupons, cleaning coupons, environmental witness plates and reference materials from manufacturing facilities. Purity and cleanliness of collector substrates. Semi-conductor vendors provided surface cleanliness data and some purity data. Purity for specific elements of interest was verified by science team members in their laboratories [4]. Curation archived procurement and shipping records, analysis reports, and non-proprietary fabrication data. A physical archive of flight collector reference materials is maintained for future use so additional data can be collected as analytical techniques improve. These are of increased value due to the hard landing upon re-entry. Cleaning and cleanliness assessments of flight hardware. Cleaning of the science canister payload was performed at JSC in a dedicated ISO 4 cleanroom using ultrapure water (UPW). The cleanliness of this UPW was monitored throughout processing. The archive for the clean lab also includes airborne particle counts, airborne molecular and inorganic contamination measurements as well as cleanroom construction material coupons and witness coupons. Hardware cleanliness was assessed by particle counts in rinse water batches. This information is recorded in batch cleaning forms and logbooks, and are, perhaps, of decreased value due to the hard landing. Post-flight: Curation-generated data. The curation handling history of each Genesis sample is documented in a typical astromaterials sample database which captures sample location, physical description and characterization data. Samples have a “shelf life”. Crucial to the preservation of samples is ongoing documentation of the sample environment, initially under curatorial control but is now a separate facility function with requires coordination. PI-generated data. Data on sample characterization and cleaning techniques continues to be generated by sample users [5]. These are often captured in LPSC abstracts, but these “engineering” results often are not publishable as stand-alone papers. We are actively looking for ways to make this information more accessible to users. Ion implants into samples have aided science return and can be shared among investigators. These (and similar) materials should be added to the curatorial collection with appropriate process and characterization data generated externally. Summary: Complete data archives for returned astromaterial samples must be broad in types and formats, and inclusive of environmental monitoring.

Genesis↗

FY22 Grid Modernization & Energy Storage Program: Accomplishments & Impacts

Sandia’s Grid Modernization and Energy Storage program works to advance a national vision of a secure, resilient, and sustainable electric system for all users. Our achievements reflect a strategic approach combining technology development; modeling, simulation, and data analytics; and partnered demonstrations and outreach to further the adoption of advanced grid and storage technologies. Our FY22 efforts leverage the strengths of our partnerships—spanning Sandia’s core science and technology competencies as well as external technology leaders—to develop the solutions today which enable the grid of tomorrow. Much of the material in this report comes from the separate 2022 Accomplishments Report compiled by our Energy Storage subprogram team, a cornerstone of our grid research and achievements. The Grid Energy Storage Program at Sandia is focused on making energy storage cost-effective through research and development (R&D) in new battery technologies, advanced power electronics and power conversion systems, improved safety and reliability for energy storage systems, analytical tools for the valuation of energy storage, and the validation of new energy storage technologies through demonstration projects. During the 2022 fiscal year, Sandia executed R&D work supported by the U.S. Department of Energy’s (DOE) Office of Electricity – Energy Storage Program under the leadership of Dr. Imre Gyuk. This report indicates key areas of research and engagement and summarizes the impact of Sandia’s contributions through notable accomplishments, journal publications, patents, and technical conferences and presentations. It is provided with the hope that readers discover ways we can further team to create our modern grid and apply the outcomes of our efforts. The bulk of work described herein is funded by the DOE Office of Electricity and key programs within the DOE Office of Energy Efficiency and Renewable Energy. As we indicated in our report from last year, the contributors to our successes are too numerous to name here, though our team wishes to express our deep gratitude to the numerous program and project sponsors at the US Department of Energy, who often function equally as technical collaborators; our many partners in industry, academia, utilities, and other national labs; and fellow researchers and business partners at Sandia whose leadership and creativity have enabled the accomplishments described herein.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

MethodOpt: a Shiny-based graphical user interface for multivariate optimization of sampling and analytical instrumentation

Method optimization is an important step in producing useful data in various experimental settings involving the use of sampling and analytical instrumentation, such as gas-chromatography mass-spectrometry or other analytical techniques. However, traditional optimization techniques often lack the sophistication of more modern optimization techniques developed in areas of applied mathematics. A graphical user interface has been developed that implements a multivariate, multi-objective optimization technique for spectra-generating sampling and analytical instrumentation, which saves substantial time and resources compared to the more traditional approaches to method development.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Data Archive and Portal (DAP) Platform for Solid Phase Processing Technologies

The scale and speed of data generated by modern scientific experiments have constantly challenged the research community to store, curate, manage and optimally use it to drive scientific discoveries. In this work, we have developed a data archive and portal (DAP) platform including analytics capabilities to collect, curate, and manage data and metadata stream for solid phase processing (SPP) techniques. We successfully hosted around ~347K files of data related to processing parameters, microscopic images, and spectroscopic data related to solid phase processing. The DAP platform for SPP will establish an enduring capability to support machine learning and grow collaboration at the intersection of materials science and data science.

36 MATERIALS SCIENCE↗

Modeling of Transient Flow Mixing of Streams Injected into a Mixing Chamber

Ignition is recognized as one the critical drivers in the reliability of multiple-start rocket engines. Residual combustion products from previous engine operation can condense on valves and related structures thereby creating difficulties for subsequent starting procedures. Alternative ignition methods that require fewer valves can mitigate the valve reliability problem, but require improved understanding of the spatial and temporal propellant distribution in the pre-ignition chamber. Current design tools based mainly on one-dimensional analysis and empirical models cannot predict local details of the injection and ignition processes. The goal of this work is to evaluate the capability of the modern computational fluid dynamics (CFD) tools in predicting the transient flow mixing in pre-ignition environment by comparing the results with the experimental data. This study is a part of a program to improve analytical methods and methodologies to analyze reliability and durability of combustion devices. In the present paper we describe a series of detailed computational simulations of the unsteady mixing events as the cold propellants are first introduced into the chamber as a first step in providing this necessary environmental description. The present computational modeling represents a complement to parallel experimental simulations' and includes comparisons with experimental results from that effort. A large number of rocket engine ignition studies has been previously reported. Here we limit our discussion to the work discussed in Refs. 2, 3 and 4 which is both similar to and different from the present approach. The similarities arise from the fact that both efforts involve detailed experimental/computational simulations of the ignition problem. The differences arise from the underlying philosophy of the two endeavors. The approach in Refs. 2 to 4 is a classical ignition study in which the focus is on the response of a propellant mixture to an ignition source, with emphasis on the level of energy needed for ignition and the ensuing flame propagation issues. Our focus in the present paper is on identifying the unsteady mixing processes that provide the propellant mixture in which the ignition source is to be placed. In particular, we wish to characterize the spatial and temporal mixture distribution with a view toward identifying preferred spatial and temporal locations for the ignition source. As such, the present work is limited to cold flow (pre-ignition) conditions

Voytovych, Dmytro M.↗

Estimating Cosmological Constraints from Galaxy Cluster Abundance using Simulation-Based Inference

Inferring the values and uncertainties of cosmological parameters in a cosmology model is of paramount importance for modern cosmic observations. In this paper, we use the simulation-based inference (SBI) approach to estimate cosmological constraints from a simplified galaxy cluster observation analysis. Using data generated from the Quijote simulation suite and analytical models, we train a machine learning algorithm to learn the probability function between cosmological parameters and the possible galaxy cluster observables. The posterior distribution of the cosmological parameters at a given observation is then obtained by sampling the predictions from the trained algorithm. Our results show that the SBI method can successfully recover the truth values of the cosmological parameters within the 2σ limit for this simplified galaxy cluster analysis, and acquires similar posterior constraints obtained with a likelihood-based Markov Chain Monte Carlo method, the current state-of the-art method used in similar cosmological studies.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

High Resolution Nature Runs and the Big Data Challenge

NASA's Global Modeling and Assimilation Office at Goddard Space Flight Center is undertaking a series of very computationally intensive Nature Runs and a downscaled reanalysis. The nature runs use the GEOS-5 as an Atmospheric General Circulation Model (AGCM) while the reanalysis uses the GEOS-5 in Data Assimilation mode. This paper will present computational challenges from three runs, two of which are AGCM and one is downscaled reanalysis using the full DAS. The nature runs will be completed at two surface grid resolutions, 7 and 3 kilometers and 72 vertical levels. The 7 km run spanned 2 years (2005-2006) and produced 4 PB of data while the 3 km run will span one year and generate 4 BP of data. The downscaled reanalysis (MERRA-II Modern-Era Reanalysis for Research and Applications) will cover 15 years and generate 1 PB of data. Our efforts to address the big data challenges of climate science, we are moving toward a notion of Climate Analytics-as-a-Service (CAaaS), a specialization of the concept of business process-as-a-service that is an evolving extension of IaaS, PaaS, and SaaS enabled by cloud computing. In this presentation, we will describe two projects that demonstrate this shift. MERRA Analytic Services (MERRA/AS) is an example of cloud-enabled CAaaS. MERRA/AS enables MapReduce analytics over MERRA reanalysis data collection by bringing together the high-performance computing, scalable data management, and a domain-specific climate data services API. NASA's High-Performance Science Cloud (HPSC) is an example of the type of compute-storage fabric required to support CAaaS. The HPSC comprises a high speed Infinib and network, high performance file systems and object storage, and a virtual system environments specific for data intensive, science applications. These technologies are providing a new tier in the data and analytic services stack that helps connect earthbound, enterprise-level data and computational resources to new customers and new mobility-driven applications and modes of work. In our experience, CAaaS lowers the barriers and risk to organizational change, fosters innovation and experimentation, and provides the agility required to meet our customers' increasing and changing needs

big data analysis↗