Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “big data applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

172 records · Page 10

FIU Project 3: Waste and D&D Engineering and Technology Development [Slides]

This project focuses on delivering solutions under deactivation and decommissioning (D&D) in support of DOE EM-4.11 as well as IT development for environmental applications (KM-IT for EM-4.11) and waste & material management (WIMS for EM-4.22). All technology development related activities will also engage the Office of Technology Development (EM-3.2). This work is also relevant to infrastructure management activities being carried out at DOE sites such as Oak Ridge, Savannah River, Hanford, Idaho and Portsmouth. As appropriate and within the parameters of the DOE-FIU Cooperative Agreement (CA), coordination at the proper level will occur with the sites and national laboratories involved in the project research efforts as well as with the points-of-contact at DOE HQ (e.g., HQ Project Leads, HQ Field Liaisons, Office of Technology Development, CA Technical Monitor, COR, etc.).

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Orchestration of materials science workflows for heterogeneous resources at large scale

In the era of big data, materials science workflows need to handle large-scale data distribution, storage, and computation. Any of these areas can become a performance bottleneck. We present a framework for analyzing internal material structures (e.g., cracks) to mitigate these bottlenecks. We demonstrate the effectiveness of our framework for a workflow performing synchrotron X-ray computed tomography reconstruction and segmentation of a silica-based structure. Our framework provides a cloud-based, cutting-edge solution to challenges such as growing intermediate and output data and heavy resource demands during image reconstruction and segmentation. Specifically, our framework efficiently manages data storage, scaling up compute resources on the cloud. The multi-layer software structure of our framework includes three layers. A top layer uses Jupyter notebooks and serves as the user interface. A middle layer uses Ansible for resource deployment and managing the execution environment. A low layer is dedicated to resource management and provides resource management and job scheduling on heterogeneous nodes (i.e., GPU and CPU). At the core of this layer, Kubernetes supports resource management, and Dask enables large-scale job scheduling for heterogeneous resources. The broader impact of our work is four-fold: through our framework, we hide the complexity of the cloud’s software stack to the user who otherwise is required to have expertise in cloud technologies; we manage job scheduling efficiently and in a scalable manner; we enable resource elasticity and workflow orchestration at a large scale; and we facilitate moving the study of nonporous structures, which has wide applications in engineering and scientific fields, to the cloud. While we demonstrate the capability of our framework for a specific materials science application, it can be adapted for other applications and domains because of its modular, multi-layer architecture.

97 MATHEMATICS AND COMPUTING↗

Performing Bayesian Analyses With AZURE2 Using BRICK: An Application to the 7 Be System

Phenomenological R-matrix has been a standard framework for the evaluation of resolved resonance cross section data in nuclear physics for many years. It is a powerful method for comparing different types of experimental nuclear data and combining the results of many different experimental measurements in order to gain a better estimation of the true underlying cross sections. Yet a practical challenge has always been the estimation of the uncertainty on both the cross sections at the energies of interest and the fit parameters, which can take the form of standard level parameters. Frequentist (χ 2 -based) estimation has been the norm. In this work, a Markov Chain Monte Carlo sampler, emcee, has been implemented for the R-matrix code AZURE2, creating the Bayesian R-matrix Inference Code Kit (BRICK). Bayesian uncertainty estimation has then been carried out for a simultaneous R-matrix fit of the 3 He (α,γ) 7 Be and 3 He (α,α) 3 He reactions in order to gain further insight into the fitting of capture and scattering data. Both data sets constrain the values of the bound state α-particle asymptotic normalization coefficients in 7 Be. The analysis highlights the need for low-energy scattering data with well-documented uncertainty information and shows how misleading results can be obtained in its absence.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

MindSynchro

This report presents the developments and results of MindSynchro project as part of DOE OE FOA 1861. DOE and Pacific Northwest National Laboratory (PNNL) have made available to FOA awardees datasets containing years of real historical data recorded from various phasor measurement units (PMUs) which are installed in three large US interconnections: Texas (IC A), Western (IC B), and Eastern (IC C). The main goal of the project, which was successfully achieved, was to develop methods for detection and identification of events which are relevant for power grid operation. Tasks performed for achieving the project goals included data exploration and pre-processing, the development and application of physics-based features, data analysis and labeling based on unsupervised learning approaches, training and testing of DSSL models for classification of events which are relevant for power grid operation, and deployment of solutions to cloud environments. The methods developed in the project can potentially provide relevant benefits to power grid asset owners/operators in general in terms of situational awareness. Two main types of outcomes can be provided by these tools: Identification of specific relevant power grid event types: Semi-supervised ML methods developed in the project can adequately employ not only the relatively scarce labeled data but also the large amount of available unlabeled data to train models for detection of specific event types. Such methods enable the application of trained models for the detection of events in a population of PMUs much larger than that associated to the labeled events. Support in data labeling / label validation: Labels are critical for training of models for identification of specific types of events. However, labeling large amounts of data is a manual and tedious process. This means that such process is error prone and is not scalable. Methods developed in the project, based on ensembles of clustering models, have been successfully employed for turning manual labeling into a scalable process. Accurate identification of specific relevant events can provide the operators with immediate situational awareness that could otherwise require hours or days of analysis from domain experts. We envision that such methods could be initially employed in support of post-mortem analysis of events and, as confidence is gained, they could be employed for online/real-time support, providing, among other benefits, insights for avoiding major events which could happen due to a combination of smaller ones. On the longer term, related methods could potentially be employed to improve protection and control.

24 POWER TRANSMISSION AND DISTRIBUTION↗

iDDS: intelligent distributed dispatch and scheduling for workflow orchestration

The intelligent distributed dispatch and scheduling (iDDS) service is a versatile workflow orchestration system designed for large-scale, distributed scientific computing. iDDS extends traditional workload and data management by integrating data-aware execution, conditional logic, and programmable workflows, enabling automation of complex and dynamic processing pipelines. Originally developed for the ATLAS experiment at the large hadron collider, iDDS has evolved into an experiment-agnostic platform that supports both template-driven workflows and a Function-as-a-Task model for Python-based orchestration. This paper presents the architecture and core components of iDDS, highlighting its scalability, modular message-driven design, and integration with systems such as PanDA and Rucio. We demonstrate its versatility through real-world use cases: fine-grained tape resource optimization for ATLAS, orchestration of large Directed Acyclic Graph (DAG) workflows for the Rubin Observatory, distributed hyperparameter optimization for machine learning applications, active learning for physics analyses, and AI-assisted detector design at the electron–ion collider. By unifying workload scheduling, data movement, and adaptive decision-making, iDDS reduces operational overhead and enables reproducible, high-throughput workflows across heterogeneous infrastructures. We conclude with current challenges and future directions, including interactive, cloud-native, and serverless workflow support.

97 MATHEMATICS AND COMPUTING↗

Inverse size-dependent Stokes shift in strongly quantum confined CsPbBr 3 perovskite nanoplates

Colloidal semiconductor nanocrystals (NCs) are used as bright chromatic fluorophores for energy-efficient displays. Here we focus here on the size-dependent Stokes shift for CsPbBr 3 nanocrystals. The Stokes shift, i.e., the difference between the wavelengths of absorption and emission maxima, is crucial for display application, as it controls the degree to which light is reabsorbed by the emitting material reducing the energetic efficiency. One major impediment to the industrial adoption of NCs is that slight deviations in manufacturing conditions may result in a wide dispersion of the product's properties. A data-driven analysis of over 2000 reactions comparing two data sets, one produced via standard colloidal synthesis and the other via high-throughput automated synthesis is discussed. We show that differences in the reaction conditions of colloidal CsPbBr 3 nanocrystals yield nanocrystals with opposite Stokes shift size-dependent trends. These match the morphologies of two-dimensional nanoplatelets (NPLs) and nanocrystal cubes. The Stokes shift size dependence trend of NPLs and nanocubes is non-monotonic indicating different physics is at play for the two nanocrystal morphologies. For nanocrystals with cubic shape, with the increase of edge length, there is a significant decrease in Stokes shift values. However, for NPLs with the increase of thickness (1–4 ML), Stokes shift values will increase. The study emphasizes the transition from a spectroscopic point of view and relates the two Stokes shift trends to 2D and 0D exciton dimensionalities for the two morphologies. Our findings highlight the importance of CsPbBr 3 nanocrystal morphology for Stokes shift prediction.

36 MATERIALS SCIENCE↗

Thermopile Energy Harvesting for Subsurface Wellbore Sensors (Final Report)

Robust in situ power harvesting underlies all efforts to enable downhole autonomous sensors for real-time and long-term monitoring of CO 2 plume movement and permeance, wellbore health, and induced seismicity. This project evaluated the potential use of downhole thermopile arrays, known as thermoelectric generators (TEGs), as power sources to charge sensors for in situ real-time, long-term data capture and transmission. Real-time downhole monitoring will enable “Big Data” techniques and machine learning, using massive amounts of continuous data from embedded sensors, to quantify short- and long-term stability and safety of enhanced oil recovery and/or commercial-scale geologic CO 2 storage. This project evaluated possible placement of the TEGs at two different wellbore locations: on the outside of the casing; or on the production tubing. TEGs convert heat flux to electrical power, and in the borehole environment, would convert heat flux into or out of the borehole into power for downhole sensors. Such heat flux would be driven by pumping of cold or hot fluids into the borehole—for instance, injecting supercritical CO 2 —creating a thermal pulse that could power the downhole sensors. Hence, wireless power generation could be accomplished with in situ TEG energy harvesting. This final report summarizes the project’s efforts that accomplished the creation of a fully operational thermopile field unit, including selection of materials, laboratory benchtop experiments and thermal-hydrologic modeling for design and optimization of the field-scale power generation test unit. Finally, the report describes the field unit that has been built and presents results of performance and survivability testing. The performance and survivability testing evaluated the following: 1) downhole power generation in response to a thermal gradient produced by pumping a heated fluid down a borehole and through the field unit; and 2) component survivability and operation at elevated temperature and pressure conditions representative of field conditions. The performance and survivability testing show that TEG arrays are viable for generating ample energy to power downhole sensors, although it is important to note that developing or connecting to sensors was beyond the scope of this project. This project’s accomplishments thus traversed from a low Technical Readiness Level (TRL) on fundamental concepts of the application and modeling to TRL-5 via testing of the fully integrated field unit for power generation in relevant environments. A fully issued United States Patent covers the wellbore power harvesting technology and applications developed by this project.

47 OTHER INSTRUMENTATION↗

Opportunities for enhancing MLCommons efforts while leveraging insights from educational MLCommons earthquake benchmarks efforts

MLCommons is an effort to develop and improve the artificial intelligence (AI) ecosystem through benchmarks, public data sets, and research. It consists of members from start-ups, leading companies, academics, and non-profits from around the world. The goal is to make machine learning better for everyone. In order to increase participation by others, educational institutions provide valuable opportunities for engagement. In this article, we identify numerous insights obtained from different viewpoints as part of efforts to utilize high-performance computing (HPC) big data systems in existing education while developing and conducting science benchmarks for earthquake prediction. As this activity was conducted across multiple educational efforts, we project if and how it is possible to make such efforts available on a wider scale. This includes the integration of sophisticated benchmarks into courses and research activities at universities, exposing the students and researchers to topics that are otherwise typically not sufficiently covered in current course curricula as we witnessed from our practical experience across multiple organizations. As such, we have outlined the many lessons we learned throughout these efforts, culminating in the need for benchmark carpentry for scientists using advanced computational resources. The article also presents the analysis of an earthquake prediction code benchmark while focusing on the accuracy of the results and not only on the runtime; notedly, this benchmark was created as a result of our lessons learned. Energy traces were produced throughout these benchmarks, which are vital to analyzing the power expenditure within HPC environments. Additionally, one of the insights is that in the short time of the project with limited student availability, the activity was only possible by utilizing a benchmark runtime pipeline while developing and using software to generate jobs from the permutation of hyperparameters automatically. It integrates a templated job management framework for executing tasks and experiments based on hyperparameters while leveraging hybrid compute resources available at different institutions. The software is part of a collection called cloudmesh with its newly developed components, cloudmesh-ee (experiment executor) and cloudmesh-cc (compute coordinator).

58 GEOSCIENCES↗

Floating Island International PHASE I FINAL TECHNICAL REPORT (Advancing the Development of Floating Solar-Powered Nanobubble Aeration Systems for Use with Floating Treatment Wetlands in Natural and Man-Made Waterbodies)

More and more freshwater lakes are suffering from algae blooms, which deprive water of oxygen and lead to widespread loss of aquatic life and production of dangerous toxins. When oxygen is lacking, methane is generated in the sediments as algae is decomposed; it has been calculated that more than half of global methane emissions come from nutrient-impaired freshwater systems (Beaulieu, 2021). Nutrients, mainly from agricultural run-off, feed algae blooms and are exacerbated by warming related to climate change. Artificial aeration is frequently used to restore oxygen to a waterbody. But diffuser aeration is inefficient and expensive, needs on-site grid power, and is failing to keep pace with the demands of water as it gets warmer with climate change. In the last five years, nanobubble aeration has been introduced into freshwater applications and shows promise for rapidly increasing dissolved oxygen effectively throughout the water column. Nanobubbles deliver oxygen in bubbles that have no buoyancy, so they stay in water longer and release their oxygen more fully, compared to large bubbles that rise and burst at the surface. The prospect of this new technology answering the deficiencies of diffuser aeration drove us to initiate the current project. During our Phase I period, we have successfully tested a nanobubble aeration system that can super-saturate oxygen levels. It is operated on solar power, and both the nanobubbler and solar array are mounted on a proprietary floating island platform, that also performs biological nutrient recycling. A rudimentary system was assembled and tested on a 6.5-acre research lake at our headquarters in Montana, where it was subjected to a summer drought that reduced water levels significantly, periods of severe cold in winter, and spring conditions that included heavy rainfall and violent thunderstorms. Our twice-weekly sampling throughout the project showed that dissolved oxygen levels rose rapidly and were maintained throughout winter and spring, well into June. The nanobubbler was able to run on solar power for long periods. Methane levels were tested periodically and found to decrease as oxygen increased. The nanobubbles appeared to have no adverse impact on fish or other aquatic life exposed to them. The stability of the installation survived the weather conditions. The many problems we encountered taught us what we need to improve in the next iteration. Towards the end of our Phase I, we were awarded a supplementary state grant that enabled us to acquire a new nanobubble system for testing that runs on DC, is extremely efficient and has few moving parts. We plan to team this with low-profile solar panels, lithium batteries and a controller that provides real-time data. The ultimate goal of this project is to commercialize an affordable and scalable lake management system that will oxygenate water from the surface down to the sludge. It will enliven fisheries, prevent toxic algae blooms and – most importantly for the fate of our planet - inhibit the production and release of methane from the sediments. We anticipate that as carbon credits expand to include methane, the cost of remediating waterbodies that produce methane will be offset. We firmly believe that this is a practical technology that works and will make a big difference when fully implemented.

14 SOLAR ENERGY↗

Securing the Modern Grid: Federal Investments, Digitization, and Supply Chain Strategy

Across the United States (U.S.) grid expansion and modernization is underway, paving the way for accelerated load growth and intelligent resource management. Digitization of the grid is supported by several state and federal programs, providing support for utilities installing advanced metering infrastructure (AMI), AI-powered analytics systems, battery energy storage systems (BESS), and distributed energy resource management systems (DERMS) to transform the grid from a one-way power delivery system into an intelligent, responsive network that will enable faster load growth and power expansion of data centers for advanced artificial intelligence (AI) applications. The digital transformation of America's grid presents opportunity for increased efficiency and resiliency but also introduces new digital risks that require careful management. Digital equipment often contains several vulnerabilities such as unencrypted communication protocols, and persistent remote access capabilities that could be exploited to manipulate device settings, coordinate service disruptions, or inject false data into grid operations. These digital risks become particularly important as the grid must rapidly scale to support AI-driven data centers, which the administration has identified as essential for maintaining U.S. technological leadership and economic competitiveness. These vulnerabilities are compounded by supply chain realities: Chinese manufacturers currently produce 70-90% of essential grid components including inverters, batteries, and control systems, with the U.S. lacking domestic manufacturing capacity for critical assets like extra-high voltage transformers. Recent federal legislation has established Foreign Entity of Concern (FEOC) restrictions to address these risks, requiring projects to achieve escalating thresholds of non-FEOC content to receive tax credits while utilities work to expand sourcing channels for their supply chains and strengthen security measures. These restrictions arrive precisely when utilities face unprecedented electricity demand growth driven by the rapid growth in data centers, creating a considerable challenge: rapidly expanding infrastructure while navigating complex compliance requirements while lacking viable alternatives for many critical components. Idaho National Laboratory (INL) and its partners have developed practical approaches to help utilities navigate these intersecting challenges as they leverage federal investment to strengthen and grow the grid. These solutions include Cyber-Informed Engineering (CIE) principles that build resilience directly into systems, the Cirrus tool for secure cloud migration, and enhanced procurement guidance that embeds security requirements throughout equipment lifecycles. Federal initiatives, such as the Technical Assistance for Digital Assurance (TADA) project, provide direct support to utilities implementing these approaches while facilitating knowledge sharing across the industry. While these tools and frameworks cannot eliminate all risks inherent in foreign supply chain dependencies, they offer pragmatic pathways for strengthening security posture without sacrificing the deployment momentum essential to meeting surging electricity demand. Ultimately, securing America's digital energy infrastructure demands dedicated coordination across multiple fronts: building domestic supply chains, implementing robust digital assurance practices, and maintaining the aggressive modernization timeline necessary for reliability, resilience, and energy independence.

24 POWER TRANSMISSION AND DISTRIBUTION↗