Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “big data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

EDX ClaiMM: Digital Resources for the Critical Minerals and Materials Community

Securing critical mineral supply chains is essential for transitioning to a clean energy economy and for maintaining national security. Big-data analytics can serve as a cost-effective means of identifying new domestic critical mineral resources but only if data can be easily located and digested. Using ArcGIS Enterprise Sites, EDX ClaiMM was developed to increase the accessibility of critical minerals data, reducing time spent on data collection and integration. Hosted tools provide rapid visualization and exploration of key datasets, unlocking insights to support resource assessments.

Yesenchak, Rachel↗

Reducing Data Center Peak Cooling Demand and Energy Costs with Underground Thermal Energy Storage (UTES)

By recent estimates, data center energy demands are projected to consume between 6.7% and 12% of U.S. annual electricity generation by the year 2028, driven primarily by expanded demands from cloud services, big data analytics, and Artificial Intelligence (AI) (Shehabi et al., 2024). As much as 40% of data center total energy consumption are loads associated with the site infrastructure cooling systems, and these are often highly water consumptive (Aljbour et al., 2024). For energy system planners, this presents significant challenges to meeting and managing the anticipated loads, and especially the peak loads of projected data center deployments. Geothermal technologies offer two unique solutions to these challenges: 1) by serving loads through the deployment of new conventional and/or next-generation geothermal power technologies such as EGS and 2) through an often-overlooked opportunity to reduce data center peak cooling loads. The latter is the focus of this paper which explores Cold Underground Thermal Energy Storage ("Cold UTES") as an emerging industrial-scale geothermal cooling solution. This cooling solution is energy efficient, non-water-consumptive, and utilizes long duration energy storage (LDES) on both diurnal and seasonal time scales. Cold UTES has the potential to also function as a virtual power plant (VPP). The US Department of Energy's Geothermal Technologies Office is supporting R&D to understand the grid and system-wide value, costs, and impacts of deploying this emergent cooling solution at scale.

AI↗

Big Data Synchrophasor Monitoring and Analytics for Resiliency Tracking (BDSMART)

This report contains key findings from a project titled Big Data Synchrophasor Monitoring and Analytics for Resiliency Tracking (BDSMART), which was carried out through a collaborative effort of a team of researchers from Texas A&M Engineering Experiment Station, Temple University, and Quanta Technology, LLC. The in-kind support came from OSIsoft (acquired by AVEVA), which provided their PI Historian software to demonstrate the use case of streaming PMU data. The first section of the report describes the project goals and objectives related to the development of Machine Learning (ML) models capable of detecting and classifying events by processing phasor measurements captured in the field by Phasor Measurement Units (PMUs). The data for this study was contributed by the utilities/ISOs from the Western and Eastern interconnects and ERCOT, further referred to as Interconnect B (IC B), Interconnect A (IC A), and Interconnect C (IC C), respectively. The approach that the BDSMART Research Team proposed and the key research tasks defined by the team are outlined in this section. The next section describes the technical approach. We first discuss the data constraints related to the PMU measurements and data interpretation constraints imposed by the data contributors. They provided neither the topological information of the grid nor PMU placement locations and captured recorded data at very few locations in the system with the reporting rate of either 30 or 60 fps. The recordings are mostly positive sequence voltage, frequency, and ROCOF, and in some limited cases, three-phase voltages and currents. We then reflect on the bad data issues that stem from poor recording practices and vague definitions of the PMU status bits to supposedly be used for bad data identification. Finally, the data discovery points to imprecise time stamps with incomplete event start/end time, as well as inconsistent and incomplete event labeling, which combined make the implementation of the data models using supervising learning quite challenging. Following the data discovery study, we hypothesize that because the IC B data has the most complete labels, we should focus our model development on that data and then test it on data from other interconnects. We also define the common metrics used to evaluate the results from the ML algorithm tests. We concluded this section by summarizing the common ML models we used and explaining how we implemented and tested them. The issues from this section are expanded in the Training Dataset Report from this project. The final section of this report deals with the accomplishments and conclusions. As the accomplishments, we formulate the problem we are solving and what is achieved by solving the problem. We then reflect on each of the analytics tools we developed and point out the performance of each tool when applied to solving the mentioned problems. We reference this work for further details to the papers we published on each tool. In the conclusions, we give recommendations on how to improve future PMU recording practices to facilitate the ML algorithm implementation and guidance for the future standardization work aimed at clarifying the ambiguities associated with the PMU status bits. We finally list future tasks that can bring about further improvements in the proposed algorithms. The issues from this section are expanded in the Training, and Test Dataset Report filed at the project completion date.

97 MATHEMATICS AND COMPUTING↗

In situ feature analysis for large-scale multiphase flow simulations

The study of multiphase flow is essential for designing chemical reactors such as fluidized bed reactors (FBR), as a detailed understanding of hydrodynamics is critical for optimizing reactor performance and stability. An FBR allows scientists to conduct different types of chemical reactions involving multiphase materials, especially interaction between gas and solids. During such complex chemical processes, the formation of void regions in the reactor, generally termed as bubbles, is an important phenomenon. The study of these bubbles has a deep implication in predicting the reactor’s overall efficiency. But physical experiments needed to understand bubble dynamics are costly and non-trivial due to the technical difficulties involved and harsh working conditions of the reactors. Therefore, to study such chemical processes and bubble dynamics, a state-of-the-art computational simulation MFIX-Exa is being developed. Despite the proven accuracy of MFIX-Exa in modeling bubbling phenomena, the large-scale output data prohibits the use of traditional post hoc analysis capabilities in both storage and I/O time. Herein, to address these issues and allow the application scientists to explore the bubble dynamics in an efficient and timely manner, we have developed an end-to-end analytics pipeline that enables in situ detection of bubbles, followed by a flexible post hoc visual exploration methodology of bubble dynamics. The proposed method enables interactive analysis of bubbles, along with quantification of several bubble characteristics, enabling experts to understand the bubble interactions in detail. Positive feedback from the experts has indicated the efficacy of the proposed approach for exploring bubble dynamics in very-large-scale multiphase flow simulations.

97 MATHEMATICS AND COMPUTING↗

ENSIGN

ENSIGN is a data analytics software package offering a modern unsupervised machine learning solution for scalable discovery in Big Data. The analytics in ENSIGN are based on an advanced mathematical tool called tensor decomposition and they are optimized to run efficiently on a range of computing platforms (from small multicore Desktop platforms to large Supercomputing clusters and novel high-end memory-driven computing platforms such as HPE Superdome Flex). ENSIGN enables the user to extract deep insights from the entirety of massive-scale (100s of Gigabytes or Terabytes scale) multidimensional data. ENSIGN uncovers latent patterns in data without the user having to specify or describe what the patterns are; the user, in the first place, may not even know such patterns existed and that they have to look for such patterns. The insights gained from ENSIGN could be trailheads that can be used as starting points for deeper forensic investigation.

Baskaran, Muthu↗

Causality-respecting adaptive refinement for PINNs: enabling precise interface evolution in phase field modeling

Physics-informed neural networks (PINNs) have emerged as a powerful tool for solving physical systems described by partial differential equations (PDEs). However, their accuracy in dynamical systems, particularly those involving sharp moving boundaries with complex initial morphologies, remains a challenge. Here, this study introduces an approach combining residual-based adaptive refinement (RBAR) with causality-informed training to enhance the performance of PINNs in solving spatio-temporal PDEs. Our method employs a three-step iterative process: initial causality-based training, RBAR-guided domain refinement, and subsequent causality training on the refined mesh. Applied to the Allen-Cahn equation, a widely-used model in phase field simulations, our approach demonstrates significant improvements in solution accuracy and computational efficiency over traditional PINNs. Notably, we observe an ‘overshoot and relocate’ phenomenon in dynamic cases with complex morphologies, showcasing the method’s adaptive error correction capabilities. This synergistic interaction between RBAR and causality training enables accurate capture of interface evolution, even in challenging scenarios where traditional PINNs fail. Our framework not only resolves the limitations of uniform refinement strategies but also provides a generalizable methodology for solving a broad range of spatio-temporal PDEs. The enhanced performance of the RBAR–causality combined framework demonstrates its strong potential for advancing PINN-based modeling of physical systems characterized by complex, evolving interfaces.

Allen-Cahn equations↗

Melt pool temperature measurement and monitoring during laser powder bed fusion based additive manufacturing via single-camera two-wavelength imaging pyrometry (STWIP)

Melt pool (MP) temperature is one of the determining factors and a key signature for evaluating the properties of printed components in metal additive manufacturing (AM). The state-of-the-art measurement systems are hindered, primarily by the large-scale data acquisition and processing demands. In this work, we introduce a novel coaxial, high-speed, single-camera two-wavelength imaging pyrometer (STWIP) system as opposed to the typical utilization of multiple cameras for measuring MP temperature profiles in laser powder bed fusion (LPBF) processes. Developed on a commercial LPBF machine (EOS M290), the STWIP system demonstrated its ability to quantitatively monitor the MP temperature and its variation for 50 layers at high framerates (>30,000 fps) for a real-world application (standard fatigue specimens) print. High performance computing is employed to analyze the acquired big data (MP images), for determining each MP's average temperature and 2D temperature profile. The MP temperature evolution in the gage section of a fatigue specimen is also examined at a temporal resolution of 1 ms, by evaluating the MP temperatures in the samples' first, middle, and last layers. This report is the first of its kind on monitoring MP temperature distribution and evolution at such a large, detailed scale for longer durations in practical applications.

42 ENGINEERING↗

4th Big Data for Nuclear Power Plants Workshop 2023

The Ohio State University and Idaho National Laboratory organized the 4 th Big Data for Nuclear Power Plants Workshop in November, 2023 in Columbus, Ohio. Workshop topics were chosen to understand the challenges and gaps that need to be addressed to maximize the impact of data on the nuclear industry, as well as the associated applications and risks. Discussions were focused around six specific application areas: Operation and Maintenance; Machine Learning in Nuclear Materials and Advanced Manufacturing; Cybersecurity; High-Performance Computing and Massive Computation; Big Data and Digital Twins; and Nuclear Non-Proliferation. The opportunities, challenges, and risks identified in the six focus areas explored in this workshop are diverse, but some common themes emerge, such as the importance of data integrity, quality, coverage, privacy, and traceability. Big data and AI/ML tools can be leveraged to reduce costs, optimize human tasking, and reduce human error across various application areas. In order for the nuclear industry to benefit from big data and advanced analytic capabilities, it is essential to address challenges and risks, such as data privacy, model reliability, and computational resource availability. Learning from other industries that have successfully implemented big data and AI/ML technologies, like the aerospace industry, can help the nuclear industry successfully integrate these technologies.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Smart Mobility in the Cloud: Enabling Real-Time Situational Awareness and Cyber-Physical Control Through a Digital Twin for Traffic

This article presents the design, implementation, and use cases of the Chattanooga Digital Twin (CTwin) towards the vision for next-generation smart city applications for urban mobility management. CTwin is an end-to-end web-based platform that incorporates various aspects of the decision-making process for optimizing urban transportation systems in Chattanooga, Tennessee, to reduce traffic congestion, incidents, and vehicle fuel consumption. The platform serves as a cyberinfrastructure to collect and integrate multi-domain urban mobility data from various online repositories and Internet of Things (IoT) sensors, covering multiple urban aspects (e.g., traffic, natural hazards, weather, and safety) that are relevant to urban mobility management. The platform enables advanced capabilities for: (a) real-time situational awareness on traffic and infrastructure conditions on highways and urban roads, (b) cyber-physical control for optimizing traffic signal timing, and (c) interactive visual analytics on big urban mobility data and various metrics for traffic prediction and transportation performance evaluation. The platform is designed using a multi-level componentization paradigm and is implemented using modular and adaptive architecture, rendering it as a generalizable and extendable prototype for other urban management applications. We present several use cases to demonstrate CTwin's core capabilities for supporting decision-making in smart urban mobility management.

33 ADVANCED PROPULSION SYSTEMS↗

EDX ClaiMM

EDX ClaiMM is a centralized data & analytical platform designed to revolutionize U.S. critical minerals and materials (CMM) activities. By providing a robust digital infrastructure, ClaiMM will accelerate the combination, leveraging, and rapid utilization of vital data, advanced tools, and cutting-edge research advancements in CMM. This adaptive digital research hub connects the CMM community to essential knowledge products and offers access to interoperable datasets, databases, models, software, and tools from the National Energy Technology’s (NETL’s) Energy Data eXchange (EDX) and other authoritative sources, serving both public and private sectors. EDX ClaiMM delivers AI-informed solutions to address fundamental knowledge gaps and fosters the innovation of new techniques for enhanced characterization and recovery of CMMs within the U.S. By leveraging cloud-hosted, scalable digital infrastructure, ClaiMM meets public–private applied energy needs. It equips the CMM community with priority digital resources that harness on-site and cloud compute capabilities, enabling big data storage, advanced processing, analytics, and visualization.

Critical Materials; Critical Minerals; Rare Earth ↗

Subdiffraction Imaging of Carrier Dynamics in Halide Perovskite Semiconductors: Effects of Passivation, Morphology, and Ion Motion

In this article, we spatially resolve photocarrier dynamics in halide perovskites using time-resolved electrostatic force microscopy (trEFM) to map surface potential equilibration during photoexcitation. We present a unified interpretation of trEFM, which measures the evolution of the surface potential in response to photoexcitation. We show that trEFM measurements correlate with surface recombination velocity and carrier lifetimes, validated with time-resolved photoluminescence imaging. We further validate the interpretation of trEFM through wavelength- and intensity-dependent measurements and with drift-diffusion simulations. We compare several passivation agents, including (3-aminopropyl)trimethoxysilane (APTMS), [3-(2-aminoethylamino)propyl]trimethoxysilane (AEAPTMS), and phenethylammonium iodide (PEAI). The results reveal heterogeneity in surface potential equilibration times that correlates with perovskite film morphology and nanoscale variations in recombination dynamics following surface passivation. Not only do our results highlight the potential for further improvement of passivation strategies, but also the necessity of high spatial and temporal resolution methods, like trEFM, to evaluate next-generation semiconductors.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

HAM: Hotspot-Aware Manager for Improving Communications with 3D-Stacked Memory

merging High-Performance Computing (HPC) workloads, such as graph analytics, machine learning, and big data science, are data-intensive. Data-intensive workloads usually present fine-grained memory accesses with limited or no data locality, and thus incur frequent cache misses and low utilization of memory bandwidth. 3D-stacked memory devices such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM) can provide significantly higher bandwidth than conventional memory modules. However, the traditional interfaces and optimization methods for JEDEC DDR devices do not allow to fully exploit the potential performance of 3D-stacked memory with the massive amount of irregular memory accesses of data-intensive applications. In this paper, we propose a novel Hotspot-Aware Manager (HAM) infrastructure for 3D-stacked memory devices capable of optimizing memory access streams via request aggregation, hotspot detection, and in-memory prefetching. %and an associated hotspot-aware page policy. We present the HAM design and implementation, and simulate it on a system using RISC-V embedded cores with attached HMC devices. We extensively evaluate HAM with over 12 benchmarks and applications representing diverse irregular memory access patterns. The results show that, on average, HAM reduces redundant requests by 37.51\% and increases the prefetch buffer hit rate by 4.2 times, compared to a baseline streaming prefetcher. On the selected benchmark set, HAM provides performance gains of 21.81\% in average (up to 34.28\%) and power savings of 35.07\% over a standard 3D-stacked memory.

Wang, Xi↗

A Digital Twin Framework Utilizing Machine Learning for Robust Predictive Maintenance: Enhancing Tire Health Monitoring

We introduce a novel digital twin (DT) framework for the predictive maintenance of long-term physical systems. Using monitoring tire health as an application, we show how the DT framework can be used to enhance automotive safety and efficiency, and how the technical challenges can be overcome using a three-step approach. First, to manage the data complexity over a long operation span, we employ data reduction techniques to concisely represent physical tires using historical performance and usage data. Relying on these data, for fast real-time prediction, we train a transformer-based model offline on our concise dataset to predict future tire health over time, represented as remaining casing potential (RCP). Based on our architecture, our model quantifies both epistemic and aleatoric uncertainties, providing reliable confidence intervals around predicted RCP. Second, to incorporate real-time data, we update the predictive model in the DT framework, ensuring its accuracy throughout its lifespan with the aid of hybrid modeling and the use of the discrepancy function. Third, to assist decision-making in predictive maintenance, we implement a tire state decision algorithm, which strategically determines the optimal timing for tire replacement based on RCP forecasted by our transformer model. This approach ensures that our DT accurately predicts system health, continually refines its digital representation, and supports predictive maintenance decisions. Furthermore, our framework effectively embodies a physical system, leveraging big data and machine learning (ML) for predictive maintenance, model updates, and decision-making.

advanced computing infrastructure↗

DICER: Data Intensive Computing Environment and Runtime for Evaluating Unprecedented Scale of Geospatial-Temporal Human Mobility Data

With the significant increase in sources and volume of human mobility data through commercial data vendors as well as microsimulation of cities, the scale of geospatial-temporal data to analyze and assess for mobility characterization has grown to the level of Big Data. There are mobility related commercial organizations deploying scalable computing, but often the system architecture, workflow, and intermediate processing components are not fully disclosed in relevant scope. Current research literature has a notable lack of studies demonstrating architectures and workflows for human mobility analytics that are implemented on a TeraByte scale of geospatial-temporal data. In this context, this paper presents a hyperscale-level system solution named DICER (Data Intensive Computing Environment and Runtime) for processing and analytics of geospatial-temporal data at big data scale. Although the cluster computing architecture of DICER with Apache Spark job running on Kubernetes cluster is not new, there are innovations in the workflow, hierarchical processing logic, and a wide range of intermediate preprocessing and mobility metrics calculation. We have performed case studies to validate the effectiveness of DICER system solution by performing detailed analytics and assessment of human mobility microsimulation output at three different scopes and scale, including a usecase with 16.97 TeraByte and 259.2 Billion rows of data. In addition, we have presented another case study of utilizing DICER to perform the same mobility processing and comparative analytics on large-scale commercially available geospatial-temporal data. All these case studies validate the efficiency and usefulness of DICER in computing population mobility characteristics from geospatial-temporal trajectory data at an unprecedented scale (not only just data volume, but also combination of: number of user entities, temporal frequency, spatial resolution, data duration).

De, Debraj↗

An overview of visualization and visual analytics applications in water resources management

Recent advances in information, communication, and environmental monitoring technologies have increased the availability, spatiotemporal resolution, and quality of water-related data, thereby leading to the emergence of many innovative big data applications. Among these applications, visualization and visual analytics, also known as the visual computing techniques, empower the synergy of computational methods (e.g., machine learning and statistical models) with human reasoning to improve the understanding and solution toward complex science and engineering problems. These approaches are frequently integrated with geographic information systems and cyberinfrastructure to provide new opportunities and methods for enhancing water resources management. Here, we present a comprehensive review of recent hydroinformatics applications that employ visual computing techniques to (1) support complex data-driven research problems, and (2) support the communication and decision-makings in the water resources management sector. Then, we conduct a technical review of the state-of-the-art web-based visualization technologies and libraries to share our experiences on developing shareable, adaptive, and interactive visualizations and visual interfaces for water resources management applications. We close with a vision that applies the emerging visual computing technologies and paradigms to develop the next generation of hydroinformatics applications.

54 ENVIRONMENTAL SCIENCES↗