Power system control rooms of the future: integrated big data analytics for security and resilience
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Abstract There is an emerging interest for tensor factorization applications in big‐data analytics and machine learning. To speed up the factorization of extra‐large datasets, organized in multidimensional arrays (also known as tensors), easy to compute compression‐based tensor representations, such as, Tucker and tensor train formats, are used to approximate the initial large‐tensor. Further, tensor factorization is used to extract latent features that can facilitate discoveries of new mechanisms and signatures hidden in the data, where the explainability of the latent features is of principal importance. Nonnegative tensor factorization extracts latent features that are naturally sparse and parts of the data, which makes them easily interpretable. However, to take into account available domain knowledge and subject matter expertise, often additional constraints need to be imposed, which lead us to canonical decomposition with linear constraints (CANDELINC), a canonical polyadic decomposition with rank deficient factors. In CANDELINC, Tucker compression is used as a preprocessing step, which lead to a larger residual error but to more explainable latent features. Here, we propose a nonnegative CANDELINC (nnCANDELINC) accomplished via a specific nonnegative Tucker decomposition; we refer to as minimal or canonical nonnegative Tucker. We derive several results required to understand the specificity of nnCANDELINC, focusing on the difficulties of preserving the nonnegative rank of a tensor to its Tucker core and comparing the real valued to nonnegative case. Finally, we demonstrate nnCANDELINC performance on synthetic and real‐world examples.
With the recent surge in big data analytics for hyperdimensional data, there is a renewed interest in dimensionality reduction techniques. In order for these methods to improve performance gains and understanding of the underlying data, a proper metric needs to be identified. This step is often overlooked, and metrics are typically chosen without consideration of the underlying geometry of the data. Here, in this paper, we present a method for incorporating elastic metrics into the t-distributed stochastic neighbour embedding (t-SNE) and Uniform Manifold Approximation and Projection (UMAP). We apply our method to functional data, which is uniquely characterized by rotations, parameterization and scale. If these properties are ignored, they can lead to incorrect analysis and poor classification performance. Through our method, we demonstrate improved performance on shape identification tasks for three benchmark data sets (MPEG-7, Car data set and Plane data set of Thankoor), where we achieve 0.77, 0.95 and 1.00 F1 score, respectively.
Interconnectivity has become a substratum of technology as the benefits of data-driven functionality are being realized in nearly all industries. Increased connectivity of Operational Technology (OT) exacerbates cyber risks because Industrial Control Systems (ICS) are becoming exposed to the Internet. These exposures are often done inadvertently through misconfigurations as additional network devices come online. Attack surface management (ASM) platforms can be used to identify vulnerabilities by performing external network discovery over the Internet using web spiders. These web spiders enable big data analytics of Internet of Things (IoT) devices as identifiable information of Internet-exposed equipment are archived in searchable databases that are made publicly available. There are a multitude of ASM service providers on the market. Here, this study was conducted to evaluate several commonly known tools to determine the aggregate attack surface of control systems. Queries were crafted by targeting commonly known manufacturers and communication protocols found in OT networks. Identified devices were that categorized based on technology types. Each query was replicated between several tools to target identical ICS equipment. Findings in this paper suggested a significant variance in the exposures discovered by each tool, but unique contributions were identified for each tool when a merged attack surface was derived. Therefore, all tools should be used in aggregate.
The high-latitude carbon (C) cycle is a key feedback to the global climate system, yet because of system complexity and data limitations, there is currently disagreement over whether the region is a source or sink of C. Recent advances in big data analytics and computing power have popularized the use of machine learning (ML) algorithms to upscale site measurements of ecosystem processes, and in some cases forecast the response of these processes to climate change. Due to data limitations, however, ML model predictions of these processes are almost never validated with independent datasets. To better understand and characterize the limitations of these methods, we develop an approach to independently evaluate ML upscaling and forecasting. We mimic data-driven upscaling and forecasting efforts by applying ML algorithms to different subsets of regional process-model simulation gridcells, and then test ML performance using the remaining gridcells. In this study, we simulate C fluxes and environmental data across Alaska using ecosys, a process-rich terrestrial ecosystem model, and then apply boosted regression tree ML algorithms to training data configurations that mirror and expand upon existing AmeriFLUX eddy-covariance data availability. We first show that a ML model trained using ecosys outputs from currently-available Alaska AmeriFLUX sites incorrectly predicts that Alaska is presently a modeled net C source. Increased spatial coverage of the training dataset improves ML predictions, halving the bias when 240 modeled sites are used instead of 15. However, even this more accurate ML model incorrectly predicts Alaska C fluxes under 21st century climate change because of changes in atmospheric CO 2 , litter inputs, and vegetation composition that have impacts on C fluxes which cannot be inferred from the training data. Our results provide key insights to future C flux upscaling efforts and expose the potential for inaccurate ML upscaling and forecasting of high-latitude C cycle dynamics.
The emerging trend of the convergence of high performance computing (HPC), machine learning/deep learning (ML/DL), and big data analytics presents a host of challenges for large-scale computing campaigns that seek best practices to interleave traditional scientific simulation-based workloads with ML/DL models. A portfolio of systematic approaches to incorporate deep learning into modeling and simulation serves a vital need when we support AI for science at a computing facility. In this paper, we evaluate several strategies for deploying deep learning surrogate models in a representative physics application on supercomputers at the Oak Ridge Leadership Computing Facility (OLCF). We discuss a set of recommended deployment architectures and implementation approaches. We analyze and evaluate these alternatives and show their performance and scalability up to 1000 GPUs on two mainstream platforms equipped with different deep learning hardware and software stacks.
The convergence of edge computing, big data analytics, and AI with traditional scientific calculations is increasingly being adopted in HPC workflows. Workflow management systems are crucial for managing and orchestrating these complex computational tasks. However, it is difficult to identify patterns within the growing population of HPC workflows. Serverless has emerged as a novel computing paradigm, offering dynamic resource allocation, quick response time, fine-grained resource management and auto-scaling. In this paper, we propose a framework to enable HPC scientific workflows on serverless. Our approach integrates a widely used traditional HPC workflow generator with an HPC serverless workflow management system to create benchmark suites of scientific workflows with diverse characteristics. These workflows can be executed on different serverless platforms. We comprehensively compare executing workflows on traditional local containers and serverless computing platforms. Our results show that serverless can reduce CPU and memory usage respectively by 78.11% and 73.92% without compromising performance.
Abstract MF-LOGP, a new method for determining a single component octanol–water partition coefficients ( $$LogP$$ LogP ) is presented which uses molecular formula as the only input. Octanol–water partition coefficients are useful in many applications, ranging from environmental fate and drug delivery. Currently, partition coefficients are either experimentally measured or predicted as a function of structural fragments, topological descriptors, or thermodynamic properties known or calculated from precise molecular structures. The MF-LOGP method presented here differs from classical methods as it does not require any structural information and uses molecular formula as the sole model input. MF-LOGP is therefore useful for situations in which the structure is unknown or where the use of a low dimensional, easily automatable, and computationally inexpensive calculations is required. MF-LOGP is a random forest algorithm that is trained and tested on 15,377 data points, using 10 features derived from the molecular formula to make $$LogP$$ LogP predictions. Using an independent validation set of 2713 data points, MF-LOGP was found to have an average $$RMSE$$ RMSE = 0.77 ± 0.007, $$MAE$$ MAE = 0.52 ± 0.003, and $${R}^{2}$$ R 2 = 0.83 ± 0.003. This performance fell within the spectrum of performances reported in the published literature for conventional higher dimensional models ( $$RMSE$$ RMSE = 0.42–1.54, $$MAE$$ MAE = 0.09–1.07, and $${R}^{2}$$ R 2 = 0.32–0.95). Compared with existing models, MF-LOGP requires a maximum of ten features and no structural information, thereby providing a practical and yet predictive tool. The development of MF-LOGP provides the groundwork for development of more physical prediction models leveraging big data analytical methods or complex multicomponent mixtures. Graphical Abstract
This report describes the key outcomes of research activities sponsored by the Department of Energy’s Funding Opportunity Announcement (FOA) number 1861 that was aimed at advancing the state-of-the-art in big data analytics applied to transmission-level synchrophasor measurements. The FOA resulted in eight research grants where the awardees developed machine learning and artificial intelligence tools and approaches. The commonalities in tools and approaches used by the awardees are explored, and insights gained from how the project outcomes might be operationalized are discussed. This report does not seek to comprehensively summarize all research supported by the FOA, rather it focuses on enabling the fast dissemination of major findings to the broader power systems community.
Securing critical mineral supply chains is essential for transitioning to a clean energy economy and for maintaining national security. Big-data analytics can serve as a cost-effective means of identifying new domestic critical mineral resources but only if data can be easily located and digested. Using ArcGIS Enterprise Sites, EDX ClaiMM was developed to increase the accessibility of critical minerals data, reducing time spent on data collection and integration. Hosted tools provide rapid visualization and exploration of key datasets, unlocking insights to support resource assessments.
By recent estimates, data center energy demands are projected to consume between 6.7% and 12% of U.S. annual electricity generation by the year 2028, driven primarily by expanded demands from cloud services, big data analytics, and Artificial Intelligence (AI) (Shehabi et al., 2024). As much as 40% of data center total energy consumption are loads associated with the site infrastructure cooling systems, and these are often highly water consumptive (Aljbour et al., 2024). For energy system planners, this presents significant challenges to meeting and managing the anticipated loads, and especially the peak loads of projected data center deployments. Geothermal technologies offer two unique solutions to these challenges: 1) by serving loads through the deployment of new conventional and/or next-generation geothermal power technologies such as EGS and 2) through an often-overlooked opportunity to reduce data center peak cooling loads. The latter is the focus of this paper which explores Cold Underground Thermal Energy Storage ("Cold UTES") as an emerging industrial-scale geothermal cooling solution. This cooling solution is energy efficient, non-water-consumptive, and utilizes long duration energy storage (LDES) on both diurnal and seasonal time scales. Cold UTES has the potential to also function as a virtual power plant (VPP). The US Department of Energy's Geothermal Technologies Office is supporting R&D to understand the grid and system-wide value, costs, and impacts of deploying this emergent cooling solution at scale.
This report contains key findings from a project titled Big Data Synchrophasor Monitoring and Analytics for Resiliency Tracking (BDSMART), which was carried out through a collaborative effort of a team of researchers from Texas A&M Engineering Experiment Station, Temple University, and Quanta Technology, LLC. The in-kind support came from OSIsoft (acquired by AVEVA), which provided their PI Historian software to demonstrate the use case of streaming PMU data. The first section of the report describes the project goals and objectives related to the development of Machine Learning (ML) models capable of detecting and classifying events by processing phasor measurements captured in the field by Phasor Measurement Units (PMUs). The data for this study was contributed by the utilities/ISOs from the Western and Eastern interconnects and ERCOT, further referred to as Interconnect B (IC B), Interconnect A (IC A), and Interconnect C (IC C), respectively. The approach that the BDSMART Research Team proposed and the key research tasks defined by the team are outlined in this section. The next section describes the technical approach. We first discuss the data constraints related to the PMU measurements and data interpretation constraints imposed by the data contributors. They provided neither the topological information of the grid nor PMU placement locations and captured recorded data at very few locations in the system with the reporting rate of either 30 or 60 fps. The recordings are mostly positive sequence voltage, frequency, and ROCOF, and in some limited cases, three-phase voltages and currents. We then reflect on the bad data issues that stem from poor recording practices and vague definitions of the PMU status bits to supposedly be used for bad data identification. Finally, the data discovery points to imprecise time stamps with incomplete event start/end time, as well as inconsistent and incomplete event labeling, which combined make the implementation of the data models using supervising learning quite challenging. Following the data discovery study, we hypothesize that because the IC B data has the most complete labels, we should focus our model development on that data and then test it on data from other interconnects. We also define the common metrics used to evaluate the results from the ML algorithm tests. We concluded this section by summarizing the common ML models we used and explaining how we implemented and tested them. The issues from this section are expanded in the Training Dataset Report from this project. The final section of this report deals with the accomplishments and conclusions. As the accomplishments, we formulate the problem we are solving and what is achieved by solving the problem. We then reflect on each of the analytics tools we developed and point out the performance of each tool when applied to solving the mentioned problems. We reference this work for further details to the papers we published on each tool. In the conclusions, we give recommendations on how to improve future PMU recording practices to facilitate the ML algorithm implementation and guidance for the future standardization work aimed at clarifying the ambiguities associated with the PMU status bits. We finally list future tasks that can bring about further improvements in the proposed algorithms. The issues from this section are expanded in the Training, and Test Dataset Report filed at the project completion date.
The study of multiphase flow is essential for designing chemical reactors such as fluidized bed reactors (FBR), as a detailed understanding of hydrodynamics is critical for optimizing reactor performance and stability. An FBR allows scientists to conduct different types of chemical reactions involving multiphase materials, especially interaction between gas and solids. During such complex chemical processes, the formation of void regions in the reactor, generally termed as bubbles, is an important phenomenon. The study of these bubbles has a deep implication in predicting the reactor’s overall efficiency. But physical experiments needed to understand bubble dynamics are costly and non-trivial due to the technical difficulties involved and harsh working conditions of the reactors. Therefore, to study such chemical processes and bubble dynamics, a state-of-the-art computational simulation MFIX-Exa is being developed. Despite the proven accuracy of MFIX-Exa in modeling bubbling phenomena, the large-scale output data prohibits the use of traditional post hoc analysis capabilities in both storage and I/O time. Herein, to address these issues and allow the application scientists to explore the bubble dynamics in an efficient and timely manner, we have developed an end-to-end analytics pipeline that enables in situ detection of bubbles, followed by a flexible post hoc visual exploration methodology of bubble dynamics. The proposed method enables interactive analysis of bubbles, along with quantification of several bubble characteristics, enabling experts to understand the bubble interactions in detail. Positive feedback from the experts has indicated the efficacy of the proposed approach for exploring bubble dynamics in very-large-scale multiphase flow simulations.
Physics-informed neural networks (PINNs) have emerged as a powerful tool for solving physical systems described by partial differential equations (PDEs). However, their accuracy in dynamical systems, particularly those involving sharp moving boundaries with complex initial morphologies, remains a challenge. Here, this study introduces an approach combining residual-based adaptive refinement (RBAR) with causality-informed training to enhance the performance of PINNs in solving spatio-temporal PDEs. Our method employs a three-step iterative process: initial causality-based training, RBAR-guided domain refinement, and subsequent causality training on the refined mesh. Applied to the Allen-Cahn equation, a widely-used model in phase field simulations, our approach demonstrates significant improvements in solution accuracy and computational efficiency over traditional PINNs. Notably, we observe an ‘overshoot and relocate’ phenomenon in dynamic cases with complex morphologies, showcasing the method’s adaptive error correction capabilities. This synergistic interaction between RBAR and causality training enables accurate capture of interface evolution, even in challenging scenarios where traditional PINNs fail. Our framework not only resolves the limitations of uniform refinement strategies but also provides a generalizable methodology for solving a broad range of spatio-temporal PDEs. The enhanced performance of the RBAR–causality combined framework demonstrates its strong potential for advancing PINN-based modeling of physical systems characterized by complex, evolving interfaces.
Not Available
Melt pool (MP) temperature is one of the determining factors and a key signature for evaluating the properties of printed components in metal additive manufacturing (AM). The state-of-the-art measurement systems are hindered, primarily by the large-scale data acquisition and processing demands. In this work, we introduce a novel coaxial, high-speed, single-camera two-wavelength imaging pyrometer (STWIP) system as opposed to the typical utilization of multiple cameras for measuring MP temperature profiles in laser powder bed fusion (LPBF) processes. Developed on a commercial LPBF machine (EOS M290), the STWIP system demonstrated its ability to quantitatively monitor the MP temperature and its variation for 50 layers at high framerates (>30,000 fps) for a real-world application (standard fatigue specimens) print. High performance computing is employed to analyze the acquired big data (MP images), for determining each MP's average temperature and 2D temperature profile. The MP temperature evolution in the gage section of a fatigue specimen is also examined at a temporal resolution of 1 ms, by evaluating the MP temperatures in the samples' first, middle, and last layers. This report is the first of its kind on monitoring MP temperature distribution and evolution at such a large, detailed scale for longer durations in practical applications.
Explore the source record for details and available documents.
The Ohio State University and Idaho National Laboratory organized the 4 th Big Data for Nuclear Power Plants Workshop in November, 2023 in Columbus, Ohio. Workshop topics were chosen to understand the challenges and gaps that need to be addressed to maximize the impact of data on the nuclear industry, as well as the associated applications and risks. Discussions were focused around six specific application areas: Operation and Maintenance; Machine Learning in Nuclear Materials and Advanced Manufacturing; Cybersecurity; High-Performance Computing and Massive Computation; Big Data and Digital Twins; and Nuclear Non-Proliferation. The opportunities, challenges, and risks identified in the six focus areas explored in this workshop are diverse, but some common themes emerge, such as the importance of data integrity, quality, coverage, privacy, and traceability. Big data and AI/ML tools can be leveraged to reduce costs, optimize human tasking, and reduce human error across various application areas. In order for the nuclear industry to benefit from big data and advanced analytic capabilities, it is essential to address challenges and risks, such as data privacy, model reliability, and computational resource availability. Learning from other industries that have successfully implemented big data and AI/ML technologies, like the aerospace industry, can help the nuclear industry successfully integrate these technologies.