Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Modern data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Large-Scale Trajectory Analysis via Feature Vectors

The explosion of both sensors and GPS-enabled devices has resulted in position/time data being the next big frontier for data analytics. However, many of the problems associated with large numbers of trajectories do not necessarily have an analog with many of the historic big-data applications such as text and image analysis. Modern trajectory analytics exploits much of the cutting-edge research in machine-learning, statistics, computational geometry and other disciplines. We will show that for doing trajectory analytics at scale, it is necessary to fundamentally change the way the information is represented through a feature-vector approach. We then demonstrate the ability to solve large trajectory analytics problems using this representation.

58 GEOSCIENCES↗

An analytically tractable marked power spectrum

The increasing precision of cosmology data in the modern era is calling for methods to allow the extraction of non-Gaussian information using tools beyond two-point statistics. The marked power spectrum has the potential to extract beyond two-point information in a computationally efficient way while using much of the infrastructure already available for the power spectrum. In this work we explore the marked power spectrum from an analytical perspective. In particular, we explore a low-order polynomial for the mark that allows us to better control the theoretical uncertainties and we show that with minimal new degrees of freedom the analytical results match measurements from N-body simulations for both the matter field and biased tracers in redshift space. Finally, we show that even within the limited forms of mark that we consider, there are degeneracies that can be broken by inclusion of the marked auto-spectrum or the cross-spectrum with the unmarked field. I n conclusion, we discuss future theoretical developments that would enable us to apply this approach to survey data.

79 ASTRONOMY AND ASTROPHYSICS↗

Advancing Fissile Materials Production Modeling with Adaptive Computing Environment and Simulations (ACES)

The Department of Energy’s National Nuclear Security Administration (DOE/NNSA) provides advanced capabilities to simulate the uranium enrichment process to support international negotiations on the peaceful use of nuclear energy. Uranium isotope separation centrifuges connected in a cascade configuration can produce the low-enriched uranium needed for nuclear power. However, those same centrifuges connected in a different configuration can also produce highly enriched uranium for nuclear weapons. Having the capability to assess cascade operations and identify nefarious activities promotes the peaceful uses of nuclear energy while restricting nuclear weapons proliferation. DNN R&D's Nonproliferation Stewardship Program Adaptive Computing Environment and Simulations (ACES) project is creating a modern, sustainable ecosystem of physics-based models and data-analytics tools that enables analysts to model uranium enrichment systems, simulate operational scenarios, and apply various policy options to explore potential outcomes.

07 ISOTOPE AND RADIATION SOURCES↗

Multi-system analysis of offshore geologic carbon storage: a review of open-source data science solutions

Geologic carbon storage projects are maturing worldwide and the footprint of deployment in the offshore is expanding. At present, there are ten projects in operation or that have been completed, more than 50 in construction and development, and dozens of characterization studies completed or underway. Offshore geologic carbon storage offers potential benefits over onshore geologic carbon storage. These offshore projects are generally remote in location, distant from population centers, and avoid complicated pore space rights while having abundant prospective storage potential. Some offshore fields targeted for carbon storage have comparatively fewer prior borehole penetrations except for areas that have been explored for petroleum production, minimizing potential issues such as pressure interference and infrastructure impacts. Yet offshore geologic carbon storage projects face distinctive technical and economic challenges, such as seafloor geohazards (e.g., seabed instability), expensive maritime transport, and meteorological-oceanographic conditions that can damage infrastructure and impact operations. Analytical capabilities and improved computational speeds have advanced engineering, earth and energy sciences in the wake of the arrival of modern data science over the last decade. These advancements have created an opportunity for integrated, multi-systems modeling approaches utilizing artificial intelligence and machine learning that are no longer limited by computational issues. Analytical tools developed alongside this advancement in data science can be leveraged to calibrate the potential advantages and challenges of carbon storage operations in the offshore. New methods and approaches that incorporate data science to analyze multiple aspects of engineered and natural systems can provide insights that complement the characterization and onsite engineering that traditional commercial and operational software addresses. These new methods and approaches can potentially improve the outcome of energy operations and carbon storage. Providing multi-system, science-driven data analytics enhances the knowledge base that offshore developers, operators, and regulatory bodies may draw from to improve offshore site selection and operational efficiency. Here, we provide a brief synopsis of geologic carbon storage efforts to date, an overview of the engineered and natural systems involved in offshore geologic carbon storage, and a review of publicly available, open-source, offshore and/or carbon storage related data- and science-driven tools developed by 2010 or later that are suitable for screening and assessing regions for offshore geologic carbon storage.

artificial intelligence↗

Quantifying uncertainty in analysis of shockless dynamic compression experiments on platinum. II. Bayesian model calibration

Dynamic shockless compression experiments provide the ability to explore material behavior at extreme pressures but relatively low temperatures. Typically, the data from these types of experiments are interpreted through an analytic method called Lagrangian analysis. Here, in this work, alternative analysis methods are explored using modern statistical methods. Specifically, Bayesian model calibration is applied to a new set of platinum data shocklessly compressed to 570 GPa. Several platinum equation-of-state models are evaluated, including traditional parametric forms as well as a novel non-parametric model concept. The results are compared to those in Paper I obtained by inverse Lagrangian analysis. The comparisons suggest that Bayesian calibration is not only a viable framework for precise quantification of the compression path, but also reveals insights pertaining to trade-offs surrounding model form selection, sensitivities of the relevant experimental uncertainties, and assumptions and limitations within Lagrangian analysis. The non-parametric model method, in particular, is found to give precise unbiased results and is expected to be useful over a wide range of applications. The calibration results in estimates of the platinum principal isentrope over the full range of experimental pressures to a standard error of 1.6%, which extends the results from Paper I while maintaining the high precision required for the platinum pressure standard.

Brown, Justin Lee↗

Immersive Visualization for Scientific Data Analysis

We will present the use of immersive visualization at the National Renewable Energy Laboratory (NREL), showcasing how immersive visualization is advancing scientific research and engineering practices and transforming our day-to-day operations. We are leveraging immersive visualization to support scientific discovery and engineering in various domains, including material design, computational fluid dynamics, immersive analytics, grid modernization, digital twins, and situated visualization. We have observed several benefits across four key areas: enhanced spatial judgments, improved understanding through interaction, increased capacity to embed high-dimensional data, and improved collaboration.

immersive analytics↗

A Survey of Singular Value Decomposition Methods for Distributed Tall/Skinny Data

The Singular Value Decomposition (SVD) is one of the most important matrix factorizations, enjoying a wide variety of applications across numerous application domains. In statistics and data analysis, the common applications of SVD inclue Principal Components Analysis (PCA) and regression. Usually these applications arise on data that has far more rows than columns, so-called "tall/skinny" matrices. In the big data analytics context, this may take the form of hundreds of millions to billions of rows with only a few hundred columns. There is a need, therefore, for fast, accurate, and scalable tall/skinny SVD implementations which can fully utilize modern computing resources. To that end, we present a survey of three different algorithms for computing the SVD for these kinds of tall/skinny data layouts using MPI for communication. We contextualize these with common big data analytics techniques. Finally, we present both CPU and GPU timing results from the Summit supercomputer, and discuss possible alternative approaches.

Schmidt, Drew↗

Advanced Computing, Data Science, and Artificial Intelligence Research Opportunities for Energy-Focused Transportation Science

The Energy Efficient Mobility Systems (EEMS) technology landscape is complex and rapidly evolving, which provides both tremendous opportunities and formidable challenges. Significant alterations to the mobility landscape are underway due to the advent of vehicle and infrastructure connectivity, autonomous driving, and rapid passenger- and freight-vehicle electrification. Advanced computing will play an increasingly important role in enabling the EEMS program to understand and identify the most important levers to improve the energy productivity of future integrated mobility systems. It is also driving new approaches to mobility and the research to unlock an affordable, efficient, safe, and accessible transportation future. Driving much of this change is the collection, analysis, and strategic use of massive amounts of diverse, complex data from infrastructure and vehicles with on-board sensors and data storage and transmission capabilities. Diverse and representative data are key to implementing approaches to maximize mobility energy productivity. While high-fidelity modeling of integrated transportation networks has strengthened our understanding of dynamic movement and behavior patterns, existing tools must be expanded beyond their current focus. This work necessitates data infrastructure investments (e.g., secure-streaming data platforms driven by ubiquitous sensors and video analytics) as well as investments in critical capabilities for large-scale automated analysis and organization using modern machine learning, statistics, and artificial intelligence. Other chief needs include agile, large-scale storage that can be quickly searched and queried for relevant data to support validation and model development, data-sharing agreements, and formatting standards for key data types. The future of public transit must be explored in greater detail, research must inform design, and opportunities must be identified for improving the mobility productivity of public transit in both urban and rural America.

33 ADVANCED PROPULSION SYSTEMS↗

Developing a Decision Support Engine to Enable Irrigation Modernization - Poster

Irrigation systems in the United States are some of the oldest continually utilized infrastructure in existence today, with some systems exceeding 100 years in age. They are operated to meet farming demands but are managed through a balance of varying influence: policies limiting water usage, stakeholder interests, and environmental impacts. Irrigation modernization is defined as a set of activities that update and improve existing irrigation systems, including, but not limited to improving water quantity, development of distributed energy resources for surrounding communities, ecosystem services, and improved agricultural yields. Modernizing an existing irrigation system can enable stakeholders to combat changing environmental and population demands but is difficult because the complexity involved in determining the potential benefits and consequences of irrigation modernization is high. We are combining a large amount of various geospatial, tabular, and temporal data with subject matter expertise into a decision support engine that will enable stakeholders to determine the benefits and consequences of irrigation modernization in their irrigation systems. A web-based GIS will allow the user to construct the modifications out of a palette of modernization options, which will be sent to the analytics engine for computations, and back to the web client for a graphical display and comparison of relevant metrics. Our development process involves four phases: 1) identify mechanisms of modernization, 2) identify data requirements, data streams, first principles and applicable algorithms necessary to quantify modernization mechanisms, 3) create ‘modules’ for each modernization mechanism, these modules will form the decision support engine, each capable of performing independently but can also inform other modules when needed, 4) Merge the decision support engine with a user interface, capable of ingesting user inputs and returning insights into the impacts of a modernization project as they relate to economic, environmental, monetary, and energy generation. Once complete, it is our intention that this tool will be fundamental in irrigation modernization projects, providing a strong analytical basis from which stakeholders can quickly make informed decisions regarding project development.

13 HYDRO ENERGY↗

Preventing Reverse Engineering of Critical Industrial Data with DIOD

Business analytics augmented by artificial intelligence and machine learning (AI/ML) have revolutionized the role of data in the modern world. In recent years, businesses have incorporated data into their decision-making process for better prediction, risk-assessment, content creation, etc. While such businesses often seek to leverage the full use of their data through third-party AI/ML services, they are often hampered by the risks of data leaks, reverse-engineering, stolen technology, etc. that often have disastrous consequences for businesses and their stakeholders alike. Thus, there arises a need for data masking prior to its transmission that obfuscates proprietary information while preserving the information relevant for AI/ML applications. In order to meet the needs of industrial data which are significantly different from those of data warehouses, previous work proposed an efficient time and space-scalable data masking paradigm known as the deceptive infusion of data (DIOD) methodology. The present work expands upon this work by leveraging existing reverse-engineering capabilities to facilitate the decomposition of industrial data into its proprietary and AI/ML-relevant parts, referred to as fundamental and inference metadata respectively. Both sets of metadata are further obfuscated in accordance with the DIOD methodology to create the DIOD rendition of the industrial data, which is rendered immune to reverse-engineering by discarding proprietary information and only preserving AI/ML-relevant information. Additionally, constraints of the original DIOD manuscript are relaxed using mutual information by configuring the methodology to the target AI/ML application to unlock the full potential of the DIOD methodology. As an example, data from a nuclear reactor is transformed into that from a nonlinear spring-mass system with different levels of data masking as required by the generic system and the target application.

97 MATHEMATICS AND COMPUTING↗

Selection of Global Climate Model Data for Downscaling With Generative Machine Learning and Use in the Power Planning for Alignment of Climate and Energy Systems Project

The range of results from climate models and scenarios is important to the understanding of uncertainty in power planning analysis. A U.S. Department of Energy-funded analytic project called Power Planning for Alignment of Climate and Energy Systems is developing data and analytic methods to reflect the effects of climate change on key variables for power system planning, as part of the Grid Modernization Lab Consortium. This project will select and prepare global climate model results for use in power system planning models. A related report (Evaluation of Global Climate Models for Use in Energy Analysis) assesses the performance of various global climate models from the Coupled Model Intercomparison Project Phase 6 data archive for their historical skill with respect to energy system performance and for their future projections under multiple climate change scenarios. Building from that report, we describe the selection of a climate scenario (Shared Socioeconomic Pathway [SSP] 2-4.5) and five climate models: TaiESM1, EC-Earth3-CC, GFDL-CM4, EC-Earth3-Veg, and MPI-ESM1-2-HR. We describe the model selection criteria, which were based on the quality of the match between model results under historical conditions and on the representation of the range of future values for several variables. These results will be downscaled via an open-source generative machine learning method called Super-Resolution for Renewable Energy Resource Data with Climate Change Impacts.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Materials Engineering of Violin Soundboards by Stradivari and Guarneri

Abstract We investigated the material properties of Cremonese soundboards using a wide range of spectroscopic, microscopic, and chemical techniques. We found similar types of spruce in Cremonese soundboards as in modern instruments, but Cremonese spruces exhibit unnatural elemental compositions and oxidation patterns that suggest artificial manipulation. Combining analytical data and historical information, we may deduce the minerals being added and their potential functions—borax and metal sulfates for fungal suppression, table salt for moisture control, alum for molecular crosslinking, and potash or quicklime for alkaline treatment. The overall purpose may have been wood preservation or acoustic tuning. Hemicellulose fragmentation and altered cellulose nanostructures are observed in heavily treated Stradivari specimens, which show diminished second‐harmonic generation signals. Guarneri's practice of crosslinking wood fibers via aluminum coordination may also affect mechanical and acoustic properties. Our data suggest that old masters undertook materials engineering experiments to produce soundboards with unique properties.

Su, Cheng‐Kuan↗

Materials Engineering of Violin Soundboards by Stradivari and Guarneri

Abstract We investigated the material properties of Cremonese soundboards using a wide range of spectroscopic, microscopic, and chemical techniques. We found similar types of spruce in Cremonese soundboards as in modern instruments, but Cremonese spruces exhibit unnatural elemental compositions and oxidation patterns that suggest artificial manipulation. Combining analytical data and historical information, we may deduce the minerals being added and their potential functions—borax and metal sulfates for fungal suppression, table salt for moisture control, alum for molecular crosslinking, and potash or quicklime for alkaline treatment. The overall purpose may have been wood preservation or acoustic tuning. Hemicellulose fragmentation and altered cellulose nanostructures are observed in heavily treated Stradivari specimens, which show diminished second‐harmonic generation signals. Guarneri's practice of crosslinking wood fibers via aluminum coordination may also affect mechanical and acoustic properties. Our data suggest that old masters undertook materials engineering experiments to produce soundboards with unique properties.

36 MATERIALS SCIENCE↗

FY22 Grid Modernization & Energy Storage Program: Accomplishments & Impacts

Sandia’s Grid Modernization and Energy Storage program works to advance a national vision of a secure, resilient, and sustainable electric system for all users. Our achievements reflect a strategic approach combining technology development; modeling, simulation, and data analytics; and partnered demonstrations and outreach to further the adoption of advanced grid and storage technologies. Our FY22 efforts leverage the strengths of our partnerships—spanning Sandia’s core science and technology competencies as well as external technology leaders—to develop the solutions today which enable the grid of tomorrow. Much of the material in this report comes from the separate 2022 Accomplishments Report compiled by our Energy Storage subprogram team, a cornerstone of our grid research and achievements. The Grid Energy Storage Program at Sandia is focused on making energy storage cost-effective through research and development (R&D) in new battery technologies, advanced power electronics and power conversion systems, improved safety and reliability for energy storage systems, analytical tools for the valuation of energy storage, and the validation of new energy storage technologies through demonstration projects. During the 2022 fiscal year, Sandia executed R&D work supported by the U.S. Department of Energy’s (DOE) Office of Electricity – Energy Storage Program under the leadership of Dr. Imre Gyuk. This report indicates key areas of research and engagement and summarizes the impact of Sandia’s contributions through notable accomplishments, journal publications, patents, and technical conferences and presentations. It is provided with the hope that readers discover ways we can further team to create our modern grid and apply the outcomes of our efforts. The bulk of work described herein is funded by the DOE Office of Electricity and key programs within the DOE Office of Energy Efficiency and Renewable Energy. As we indicated in our report from last year, the contributors to our successes are too numerous to name here, though our team wishes to express our deep gratitude to the numerous program and project sponsors at the US Department of Energy, who often function equally as technical collaborators; our many partners in industry, academia, utilities, and other national labs; and fellow researchers and business partners at Sandia whose leadership and creativity have enabled the accomplishments described herein.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

MethodOpt: a Shiny-based graphical user interface for multivariate optimization of sampling and analytical instrumentation

Method optimization is an important step in producing useful data in various experimental settings involving the use of sampling and analytical instrumentation, such as gas-chromatography mass-spectrometry or other analytical techniques. However, traditional optimization techniques often lack the sophistication of more modern optimization techniques developed in areas of applied mathematics. A graphical user interface has been developed that implements a multivariate, multi-objective optimization technique for spectra-generating sampling and analytical instrumentation, which saves substantial time and resources compared to the more traditional approaches to method development.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Data Archive and Portal (DAP) Platform for Solid Phase Processing Technologies

The scale and speed of data generated by modern scientific experiments have constantly challenged the research community to store, curate, manage and optimally use it to drive scientific discoveries. In this work, we have developed a data archive and portal (DAP) platform including analytics capabilities to collect, curate, and manage data and metadata stream for solid phase processing (SPP) techniques. We successfully hosted around ~347K files of data related to processing parameters, microscopic images, and spectroscopic data related to solid phase processing. The DAP platform for SPP will establish an enduring capability to support machine learning and grow collaboration at the intersection of materials science and data science.

36 MATERIALS SCIENCE↗

Estimating Cosmological Constraints from Galaxy Cluster Abundance using Simulation-Based Inference

Inferring the values and uncertainties of cosmological parameters in a cosmology model is of paramount importance for modern cosmic observations. In this paper, we use the simulation-based inference (SBI) approach to estimate cosmological constraints from a simplified galaxy cluster observation analysis. Using data generated from the Quijote simulation suite and analytical models, we train a machine learning algorithm to learn the probability function between cosmological parameters and the possible galaxy cluster observables. The posterior distribution of the cosmological parameters at a given observation is then obtained by sampling the predictions from the trained algorithm. Our results show that the SBI method can successfully recover the truth values of the cosmological parameters within the 2σ limit for this simplified galaxy cluster analysis, and acquires similar posterior constraints obtained with a likelihood-based Markov Chain Monte Carlo method, the current state-of the-art method used in similar cosmological studies.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗