Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “big data applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Augmentation of WRF-Hydro to simulate overland-flow- and streamflow-generated debris flow susceptibility in burn scars

In steep wildfire-burned terrains, intense rainfall can produce large runoff that can trigger highly destructive debris flows. However, the ability to accurately characterize and forecast debris flow susceptibility in burned terrains using physics-based tools remains limited. Here, we augment the Weather Research and Forecasting Hydrological modeling system (WRF-Hydro) to simulate both overland and channelized flows and assess postfire debris flow susceptibility over a regional domain. We perform hindcast simulations using high-resolution weather-radar-derived precipitation and reanalysis data to drive non-burned baseline and burn scar sensitivity experiments. Our simulations focus on January 2021 when an atmospheric river triggered numerous debris flows within a wildfire burn scar in Big Sur – one of which destroyed California's famous Highway 1. Compared to the baseline, our burn scar simulation yields dramatic increases in total and peak discharge and shorter lags between rainfall onset and peak discharge, consistent with streamflow observations at nearby US Geological Survey (USGS) streamflow gage sites. For the 404 catchments located in the simulated burn scar area, median catchment-area-normalized peak discharge increases by ~ 450 % compared to the baseline. Catchments with anomalously high catchment-area-normalized peak discharge correspond well with post-event field-based and remotely sensed debris flow observations. We suggest that our regional postfire debris flow susceptibility analysis demonstrates WRF-Hydro as a compelling new physics-based tool whose utility could be further extended via coupling to sediment erosion and transport models and/or ensemble-based operational weather forecasts. Given the high-fidelity performance of our augmented version of WRF-Hydro, as well as its potential usage in probabilistic hazard forecasts, we argue for its continued development and application in postfire hydrologic and natural hazard assessments.

54 ENVIRONMENTAL SCIENCES↗

T-FSM: A Scalable Distributed Task-Based System for Frequent Subgraph Pattern Mining from a Big Graph

Finding frequent subgraph patterns in a big graph is an important problem with many applications such as classifying chemical compounds and building indexes to speed up graph queries. Since this problem is NP-hard, some recent parallel and distributed systems have been developed to accelerate the mining. However, they often have a huge memory cost, very long running time, suboptimal load balancing, poor scale-out capability, and possibly inaccurate results. In this article, we propose an efficient system called T-FSM for parallel mining of frequent subgraph patterns in a big graph. T-FSM supports a new anti-monotonic frequentness measure called Fraction-Score, which is more accurate than the widely used MNI measure. The execution engine of T-FSM supports both intra-machine parallelism and inter-machine parallelism. For intra-machine parallelism, T-FSM adopts a novel task-based execution model to ensure high multithreading concurrency, bounded memory consumption, and effective load balancing. For inter-machine parallelism, T-FSM ensures good scale-out performance with a lightweight pattern rebalancing approach that reduces workload skewness of pattern evaluations among machines. To avoid recomputing the contexts for migrated patterns, we design a novel context cache table to support concurrent and asynchronous requesting and caching of remote context data, which can timely evict and garbage collect used pattern contexts that are no longer needed to keep memory consumption bounded. Extensive experiments show that T-FSM is orders of magnitude faster than existing state-of-the-art parallel systems (more than 10×, 51×, 131×, 55× speedup over ScaleMine, DistGraph, Pangolin and Peregrine, respectively) and distributed systems (more than 42× and 88× over ScaleMine and DistGraph, respectively) for frequent subgraph pattern mining, and it scales out satisfactorily to 512 CPU cores on the Polaris supercomputer at Argonne National Laboratory.

97 MATHEMATICS AND COMPUTING↗

Multiresolution With Super-Compact Wavelets

The solution data computed from large scale simulations are sometimes too big for main memory, for local disks, and possibly even for a remote storage disk, creating tremendous processing time as well as technical difficulties in analyzing the data. The excessive storage demands a corresponding huge penalty in I/O time, rendering time and transmission time between different computer systems. In this paper, a multiresolution scheme is proposed to compress field simulation or experimental data without much loss of important information in the representation. Originally, the wavelet based multiresolution scheme was introduced in image processing, for the purposes of data compression and feature extraction. Unlike photographic image data which has rather simple settings, computational field simulation data needs more careful treatment in applying the multiresolution technique. While the image data sits on a regular spaced grid, the simulation data usually resides on a structured curvilinear grid or unstructured grid. In addition to the irregularity in grid spacing, the other difficulty is that the solutions consist of vectors instead of scalar values. The data characteristics demand more restrictive conditions. In general, the photographic images have very little inherent smoothness with discontinuities almost everywhere. On the other hand, the numerical solutions have smoothness almost everywhere and discontinuities in local areas (shock, vortices, and shear layers). The wavelet bases should be amenable to the solution of the problem at hand and applicable to constraints such as numerical accuracy and boundary conditions. In choosing a suitable wavelet basis for simulation data among a variety of wavelet families, the supercompact wavelets designed by Beam and Warming provide one of the most effective multiresolution schemes. Supercompact multi-wavelets retain the compactness of Haar wavelets, are piecewise polynomial and orthogonal, and can have arbitrary order of approximation. The advantages of the multiresolution algorithm are that no special treatment is required at the boundaries of the interval, and that the application to functions which are only piecewise continuous (internal boundaries) can be efficiently implemented. In this presentation, Beam's supercompact wavelets are generalized to higher dimensions using multidimensional scaling and wavelet functions rather than alternating the directions as in the 1D version. As a demonstration of actual 3D data compression, supercompact wavelet transforms are applied to a 3D data set for wing tip vortex flow solutions (2.5 million grid points). It is shown that high data compression ratio can be achieved (around 50:1 ratio) in both vector and scalar data set.

Lee, Dohyung↗

Marketing Remote Sensing Data for North Pacific Fisheries Development and Management

Fish poaching, drug trafficking, ocean dumping, and other illegal activities are important problems on the high seas and in national economic zones. The primary thrust of the EOCAP II project, "Marketing Remote Sensing Data for North Pacific Fisheries Development and Management", was to use space-based sensors to improve the effectiveness of marine monitoring, control, and surveillance (MCS). Our initial objectives were to concentrate on the development of MCS tools using Advanced Very High Resolution Radiometry (AVHRR) and Synthetic Aperture Radar (SAR) data. Although we have successfully completed development of an initial version of our SAR-based monitoring tool (OmniVision), project activity has resulted in a much broader application of space-based assets to marine applications. Based in part on work commenced within EOCAP II, a new company, Ocean and Coastal Environmental Sensing, Inc. (OCENS), has been launched and the development of several new software products outside of the MCS arena initiated. One of those products, SeaStation, is near completion with a Fall, 1995 release date. Equity investment in OCENS now totals $70,000-with an additional amount being sought in the first round of financing. One of the pre-eminent objectives of EOCAP II is to make contributions to the US economy and job growth through the expansion of commercial uses of remotely sensed data. OCENS and the software products it is introducing into marine and coastal zone markets responds to this primary object*e. EOCAP II funding leveraged the market and technical know-how of OCENS founders into smart products that benefit marine and coastal zone users. Although technical difficulties and geopolitical shifts damaged the commercial feasibility of initial project objectives, the flexibility of the EOCAP II program now permits long-term business success. This in no small part stems from the fact that the EOCAP program recognizes the realities of small and start-up businesses and does not attempt to force these conditions to fit the apparent needs of big government. Instead, EOCAP works with those who know their market best in order to produce successful products and expanding businesses.

Source record↗

Spinoff 2005

Topics covered include: Lighting the Way for Quicker, Safer Healing; Discovering New Drugs on the Cellular Level; Hydrogen Sensors Boost Hybrids; Today s Models Losing Gas?; 3-D Highway in the Sky; Popping a Hole in High-Speed Pursuits; Monitoring Wake Vortices for More Efficient Airports; From Rockets to Racecars; All-Terrain Intelligent Robot Braves Battlefront to Save Lives; Keeping the Air Clean and Safe--An Anthrax Smoke Detector; Lightning Often Strikes Twice; Technology That's Ready and Able to Inspect Those Cables; Secure Networks for First Responders and Special Forces; Space Suit Spins; Cooking Dinner at Home--From the Office; Nanoscale Materials Make for Large-Scale Applications; NASA s Growing Commitment: The Space Garden; Bringing Thunder and Lightning Indoors; Forty-Year-Old Foam Springs Back With New Benefits; Experiments With Small Animals Rarely Go This Well; NASA, the Fisherman's Friend; Crystal-Clear Communication a Sweet-Sounding Success; Inertial Motion-Tracking Technology for Virtual 3-D; Then Why Do They Call Earth the Blue Planet?; Valiant 'Zero-Valent' Effort Restores Contaminated Grounds; Harnessing the Power of the Sun; Water and Air Measures That Make 'PureSense'; Remote Sensing for Farmers and Flood Watching; Pesticide-Free Device a Fatal Attraction for Mosquitoes Making the Most of Waste Energy Washing Away the Worries About Germs Celestial Software Scratches More Than the Surface A Search Engine That's Aware of Your Needs Fault-Detection Tool Has Companies 'Mining' Own Business; Software to Manage the Unmanageable; Tracking Electromagnetic Energy With SQUIDs; Taking the Risk Out of Risk Assessment; Satellite and Ground System Solutions at Your Fingertips; Structural Analysis Made 'NESSUSary'; Software of Seismic Proportions Promotes Enjoyable Learning; Making a Reliable Actuator Faster and More Affordable; Cost-Cutting Powdered Lubricant NASA s Radio Frequency Bolt Monitor: A Lifetime of Spinoffs Going End to End to Deliver High-Speed Data; Advanced Joining Technology: Simple, Strong, and Secure; Big Results From a Smaller Gearbox; Low-Pressure Generator Makes Cleanrooms Cleaner; and The Space Laser Business Model.

Source record↗

Tikiri—Towards a lightweight blockchain for IoT

Internet of Things (IoT) platforms have been deployed in several domains to enhance efficiency of business process and improve productivity. Most IoT platforms comprise of heterogeneous software and hardware components which can potentially introduce security and privacy challenges. Blockchain technology has been proposed as one of the solutions to realize IoT security by leveraging the (a) Immutable ledger, (b) Decentralized architecture and (c) Strong cryptography primitives. However, integrating blockchain platforms with IoT based applications presents several challenges due to lack of (a) acceptable performance on resource-constrained devices, (b) high transaction throughput, (c) keyword-based search and retrieve, (d) transaction back pressure operations, and (e) real-time response. In this paper, we propose a lightweight blockchain platform, “Tikiri”, for resource-constrained IoT devices. Tikiri uses Apache Kafka for the consensus and proposes new blockchain architecture to handle real-time transaction execution on the blockchain. Tikiri is characterized by functional programming and actor-based smart contract platform that realizes concurrent execution of transactions in the blockchain. Tikiri realizes a lightweight and scalable blockchain that can provides performance on the resource-constrained IoT devices.

97 MATHEMATICS AND COMPUTING↗

A method for assessing economic, environmental, and reliability tradeoffs of interregional transmission connecting ERCOT (the Texas grid) to the eastern and western grids

Reliable development of the power grid is an evolving concern for humanity due to extreme weather that frequently threatens power sector infrastructure. The state of Texas is a uniquely structured testbed for grid planners to study when looking for solutions to development, innovation, and overcoming such challenges. Because of its size and islanded structure, Texas is small enough to model, but big enough to matter. Texas is a global leader in energy production, energy consumption, and maintains an unusually diverse fuel mix. In addition, the state has experienced winter freezes, heat waves, wind storms, droughts and floods that have threatened power sector infrastructure or caused recent blackouts and calls for demand side conservation. One of the most devastating of these events was the North American winter storm, dubbed “Winter Storm Uri” by the Weather Channel, that froze the region in February 2021 and led to an extended power outage event that put the majority of Texan residents in darkness for days. While preparing to avoid such outage events in the future, various tools have been proposed to improve grid reliability, including energy efficiency, demand response, and distributed energy resources. An additional option would be to develop interregional transmission that connects the Texas grid to other national grids. To assess the merits of this idea, we developed a novel, universally-applicable and internationally-relevant framework to study how the Texas grid would evolve alongside access to various interregional ties. This method allows us to stress the synthetic grid structure and analyze how it would respond to the shock of a simulated winter storm event. Our method leverages open-source modeling tools, such as PowerGenome, pyGRETA, and GenX to synthesize unique zonal grid data, construct a consolidated network of model regions, and simulate different developmental pathways of capacity expansion and operational dispatch. We demonstrate our method with an analysis connecting the Electric Reliability Council of Texas (ERCOT), the grid that serves most of Texas, the Western Electricity Coordinating Council (WECC), the grid that serves the western half of the contiguous U.S., and the Eastern Interconnect, the grid that serves the eastern half of the contiguous U.S. Our results indicate that the cost-optimal capacity of interregional transmission connecting the ERCOT grid to other grids lies between 9–13 GW assuming baseline conditions. Building this amount of connecting capacity in one or multiple directions lowers the costs and emissions of development and operation by up to $16 billion and 257 million metric tonnes (MMT) respectively. Additionally, our results show that the interregional connections between ERCOT and other national grids reduce the amount of total load shed required through mild winter storm events. However, our results also show that there is a threshold of very extreme winter storm conditions, spanning multiple service areas, above which the connections exacerbate resource adequacy problems. Therefore, the results indicate that the connections need to be carefully planned alongside the rest of the grid infrastructure to avoid over-reliance on specific resources or technology options.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Dimensionally reduced machine learning model for predicting single component octanol–water partition coefficients

Abstract MF-LOGP, a new method for determining a single component octanol–water partition coefficients ( $$LogP$$ LogP ) is presented which uses molecular formula as the only input. Octanol–water partition coefficients are useful in many applications, ranging from environmental fate and drug delivery. Currently, partition coefficients are either experimentally measured or predicted as a function of structural fragments, topological descriptors, or thermodynamic properties known or calculated from precise molecular structures. The MF-LOGP method presented here differs from classical methods as it does not require any structural information and uses molecular formula as the sole model input. MF-LOGP is therefore useful for situations in which the structure is unknown or where the use of a low dimensional, easily automatable, and computationally inexpensive calculations is required. MF-LOGP is a random forest algorithm that is trained and tested on 15,377 data points, using 10 features derived from the molecular formula to make $$LogP$$ LogP predictions. Using an independent validation set of 2713 data points, MF-LOGP was found to have an average $$RMSE$$ RMSE = 0.77 ± 0.007, $$MAE$$ MAE = 0.52 ± 0.003, and $${R}^{2}$$ R 2 = 0.83 ± 0.003. This performance fell within the spectrum of performances reported in the published literature for conventional higher dimensional models ( $$RMSE$$ RMSE = 0.42–1.54, $$MAE$$ MAE = 0.09–1.07, and $${R}^{2}$$ R 2 = 0.32–0.95). Compared with existing models, MF-LOGP requires a maximum of ten features and no structural information, thereby providing a practical and yet predictive tool. The development of MF-LOGP provides the groundwork for development of more physical prediction models leveraging big data analytical methods or complex multicomponent mixtures. Graphical Abstract

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Breaking Barriers: Integrating Geo-Leo Aerosol Data with an Open-Source Approach

The scientific community is still examining the novel data from geostationary satellite observations and evaluating methods for effectively fusing the polar observations with various spatial and temporal resolutions. However, the merged data will present a significant ""Big Data"" challenge, including processing, storage, data discoverability, accessibility, and migration within cloud computing environments. We have developed an open-source package to fuse aerosol optical depths (AOD) products from six satellite sensors in the past four years (2019~2023), and this presentation will update our recent progress. Using this Python-based package, we produced a level 3 global (AOD) product in a quarter-degree spatial resolution every half-hour, fusing the Level 2 AOD data with the Dark Target aerosol retrieval algorithm from six satellites: three geostationary (GOES-16/17 and Himawari-8) with high temporal resolution, and three polar orbiting (TERRA/MODIS, AQUA/MODIS, and SNPP-VIIRS) with global coverage. By integrating these observations, the diurnal cycle of global AOD in this fused product can be characterized at local, regional, and global scales. Furthermore, we are committed to openness and transparency by providing our package and its associated functionalities as open-source. Our dedication to adhering to the FAIR, CARE, and TRUST principles ensures that our users can rely on the integrity and ethical standards of our work. For instance of Interoperability, this package fuses remote sensing products on demand into desired temporal and spatial domains. It can be run in a central processing unit (CPU) or a Graphics processing unit (GPU) mode. This package will empower researchers and practitioners to use satellite and sensor data efficiently in various applications and research.

Xiaohua Pan↗

Generation and representation of synthetic smart meter data

Advanced energy algorithms running at big-data scale will be necessary to identify, realize, and verify energy savings to meet government and utility goals of building energy efficiency. Any algorithm must be well characterized and validated before it is trusted to run at these scales. Smart meter data from real buildings will ultimately be required for the development, testing, and validation of these energy algorithms and processes. However, for initial development and testing, smart meter data are difficult to work with due to privacy restrictions, noise from unknown sources, data accessibility, and other concerns which can complicate algorithm development and validation. This paper describes a new methodology to generate synthetic smart meter data of electricity use in buildings using detailed building energy modeling, which aims to capture the variability and stochastics of real energy use in buildings. The methodology can create datasets tailored to represent specific scenarios with known truth and controllable amounts of synthetic noise. Knowledge of ground truth also allows the development and validation of enhanced processes which leverage building metadata, such as building type or size (floor area), in addition to smart meter data. The methodology described in this paper includes the key influencing factors of real-world building energy use including weather data, occupant-driven loads, building operation and maintenance practices, and special events. Data formats to support workflows leveraging both synthetic meter data and associated metadata are proposed and discussed. Finally, example use cases of the synthetic meter data are described to illustrate potential applications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Differentiable, Learnable, Regionalized Process-Based Models With Multiphysical Outputs can Approach State-Of-The-Art Hydrologic Prediction Accuracy

Predictions of hydrologic variables across the entire water cycle have significant value for water resources management as well as downstream applications such as ecosystem and water quality modeling. Recently, purely data-driven deep learning models like long short-term memory (LSTM) showed seemingly insurmountable performance in modeling rainfall runoff and other geoscientific variables, yet they cannot predict untrained physical variables and remain challenging to interpret. Here, we show that differentiable, learnable, process-based models (called δ models here) can approach the performance level of LSTM for the intensively observed variable (streamflow) with regionalized parameterization. We use a simple hydrologic model HBV as the backbone and use embedded neural networks, which can only be trained in a differentiable programming framework, to parameterize, enhance, or replace the process-based model's modules. Without using an ensemble or post-processor, δ models can obtain a median Nash-Sutcliffe efficiency of 0.732 for 671 basins across the USA for the Daymet forcing data set, compared to 0.748 from a state-of-the-art LSTM model with the same setup. For another forcing data set, the difference is even smaller: 0.715 versus 0.722. Meanwhile, the resulting learnable process-based models can output a full set of untrained variables, for example, soil and groundwater storage, snowpack, evapotranspiration, and baseflow, and can later be constrained by their observations. Both simulated evapotranspiration and fraction of discharge from baseflow agreed decently with alternative estimates. The general framework can work with models with various process complexity and opens up the path for learning physics from big data.

54 ENVIRONMENTAL SCIENCES↗

Applications and Techniques for Fast Machine Learning in Science

In this community review report, we discuss applications and techniques for fast machine learning (ML) in science—the concept of integrating powerful ML methods into the real-time experimental data processing loop to accelerate scientific discovery. The material for the report builds on two workshops held by the Fast ML for Science community and covers three main areas: applications for fast ML across a number of scientific domains; techniques for training and implementing performant and resource-efficient ML algorithms; and computing architectures, platforms, and technologies for deploying these algorithms. We also present overlapping challenges across the multiple scientific domains where common solutions can be found. This community report is intended to give plenty of examples and inspiration for scientific discovery through integrated and accelerated ML solutions. This is followed by a high-level overview and organization of technical advances, including an abundance of pointers to source material, which can enable these breakthroughs.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

FIU Project 3: Waste and D&D Engineering and Technology Development [Slides]

This project focuses on delivering solutions under deactivation and decommissioning (D&D) in support of DOE EM-4.11 as well as IT development for environmental applications (KM-IT for EM-4.11) and waste & material management (WIMS for EM-4.22). All technology development related activities will also engage the Office of Technology Development (EM-3.2). This work is also relevant to infrastructure management activities being carried out at DOE sites such as Oak Ridge, Savannah River, Hanford, Idaho and Portsmouth. As appropriate and within the parameters of the DOE-FIU Cooperative Agreement (CA), coordination at the proper level will occur with the sites and national laboratories involved in the project research efforts as well as with the points-of-contact at DOE HQ (e.g., HQ Project Leads, HQ Field Liaisons, Office of Technology Development, CA Technical Monitor, COR, etc.).

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Orchestration of materials science workflows for heterogeneous resources at large scale

In the era of big data, materials science workflows need to handle large-scale data distribution, storage, and computation. Any of these areas can become a performance bottleneck. We present a framework for analyzing internal material structures (e.g., cracks) to mitigate these bottlenecks. We demonstrate the effectiveness of our framework for a workflow performing synchrotron X-ray computed tomography reconstruction and segmentation of a silica-based structure. Our framework provides a cloud-based, cutting-edge solution to challenges such as growing intermediate and output data and heavy resource demands during image reconstruction and segmentation. Specifically, our framework efficiently manages data storage, scaling up compute resources on the cloud. The multi-layer software structure of our framework includes three layers. A top layer uses Jupyter notebooks and serves as the user interface. A middle layer uses Ansible for resource deployment and managing the execution environment. A low layer is dedicated to resource management and provides resource management and job scheduling on heterogeneous nodes (i.e., GPU and CPU). At the core of this layer, Kubernetes supports resource management, and Dask enables large-scale job scheduling for heterogeneous resources. The broader impact of our work is four-fold: through our framework, we hide the complexity of the cloud’s software stack to the user who otherwise is required to have expertise in cloud technologies; we manage job scheduling efficiently and in a scalable manner; we enable resource elasticity and workflow orchestration at a large scale; and we facilitate moving the study of nonporous structures, which has wide applications in engineering and scientific fields, to the cloud. While we demonstrate the capability of our framework for a specific materials science application, it can be adapted for other applications and domains because of its modular, multi-layer architecture.

97 MATHEMATICS AND COMPUTING↗

Performing Bayesian Analyses With AZURE2 Using BRICK: An Application to the 7 Be System

Phenomenological R-matrix has been a standard framework for the evaluation of resolved resonance cross section data in nuclear physics for many years. It is a powerful method for comparing different types of experimental nuclear data and combining the results of many different experimental measurements in order to gain a better estimation of the true underlying cross sections. Yet a practical challenge has always been the estimation of the uncertainty on both the cross sections at the energies of interest and the fit parameters, which can take the form of standard level parameters. Frequentist (χ 2 -based) estimation has been the norm. In this work, a Markov Chain Monte Carlo sampler, emcee, has been implemented for the R-matrix code AZURE2, creating the Bayesian R-matrix Inference Code Kit (BRICK). Bayesian uncertainty estimation has then been carried out for a simultaneous R-matrix fit of the 3 He (α,γ) 7 Be and 3 He (α,α) 3 He reactions in order to gain further insight into the fitting of capture and scattering data. Both data sets constrain the values of the bound state α-particle asymptotic normalization coefficients in 7 Be. The analysis highlights the need for low-energy scattering data with well-documented uncertainty information and shows how misleading results can be obtained in its absence.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

MindSynchro

This report presents the developments and results of MindSynchro project as part of DOE OE FOA 1861. DOE and Pacific Northwest National Laboratory (PNNL) have made available to FOA awardees datasets containing years of real historical data recorded from various phasor measurement units (PMUs) which are installed in three large US interconnections: Texas (IC A), Western (IC B), and Eastern (IC C). The main goal of the project, which was successfully achieved, was to develop methods for detection and identification of events which are relevant for power grid operation. Tasks performed for achieving the project goals included data exploration and pre-processing, the development and application of physics-based features, data analysis and labeling based on unsupervised learning approaches, training and testing of DSSL models for classification of events which are relevant for power grid operation, and deployment of solutions to cloud environments. The methods developed in the project can potentially provide relevant benefits to power grid asset owners/operators in general in terms of situational awareness. Two main types of outcomes can be provided by these tools: Identification of specific relevant power grid event types: Semi-supervised ML methods developed in the project can adequately employ not only the relatively scarce labeled data but also the large amount of available unlabeled data to train models for detection of specific event types. Such methods enable the application of trained models for the detection of events in a population of PMUs much larger than that associated to the labeled events. Support in data labeling / label validation: Labels are critical for training of models for identification of specific types of events. However, labeling large amounts of data is a manual and tedious process. This means that such process is error prone and is not scalable. Methods developed in the project, based on ensembles of clustering models, have been successfully employed for turning manual labeling into a scalable process. Accurate identification of specific relevant events can provide the operators with immediate situational awareness that could otherwise require hours or days of analysis from domain experts. We envision that such methods could be initially employed in support of post-mortem analysis of events and, as confidence is gained, they could be employed for online/real-time support, providing, among other benefits, insights for avoiding major events which could happen due to a combination of smaller ones. On the longer term, related methods could potentially be employed to improve protection and control.

24 POWER TRANSMISSION AND DISTRIBUTION↗

iDDS: intelligent distributed dispatch and scheduling for workflow orchestration

The intelligent distributed dispatch and scheduling (iDDS) service is a versatile workflow orchestration system designed for large-scale, distributed scientific computing. iDDS extends traditional workload and data management by integrating data-aware execution, conditional logic, and programmable workflows, enabling automation of complex and dynamic processing pipelines. Originally developed for the ATLAS experiment at the large hadron collider, iDDS has evolved into an experiment-agnostic platform that supports both template-driven workflows and a Function-as-a-Task model for Python-based orchestration. This paper presents the architecture and core components of iDDS, highlighting its scalability, modular message-driven design, and integration with systems such as PanDA and Rucio. We demonstrate its versatility through real-world use cases: fine-grained tape resource optimization for ATLAS, orchestration of large Directed Acyclic Graph (DAG) workflows for the Rubin Observatory, distributed hyperparameter optimization for machine learning applications, active learning for physics analyses, and AI-assisted detector design at the electron–ion collider. By unifying workload scheduling, data movement, and adaptive decision-making, iDDS reduces operational overhead and enables reproducible, high-throughput workflows across heterogeneous infrastructures. We conclude with current challenges and future directions, including interactive, cloud-native, and serverless workflow support.

97 MATHEMATICS AND COMPUTING↗

Inverse size-dependent Stokes shift in strongly quantum confined CsPbBr 3 perovskite nanoplates

Colloidal semiconductor nanocrystals (NCs) are used as bright chromatic fluorophores for energy-efficient displays. Here we focus here on the size-dependent Stokes shift for CsPbBr 3 nanocrystals. The Stokes shift, i.e., the difference between the wavelengths of absorption and emission maxima, is crucial for display application, as it controls the degree to which light is reabsorbed by the emitting material reducing the energetic efficiency. One major impediment to the industrial adoption of NCs is that slight deviations in manufacturing conditions may result in a wide dispersion of the product's properties. A data-driven analysis of over 2000 reactions comparing two data sets, one produced via standard colloidal synthesis and the other via high-throughput automated synthesis is discussed. We show that differences in the reaction conditions of colloidal CsPbBr 3 nanocrystals yield nanocrystals with opposite Stokes shift size-dependent trends. These match the morphologies of two-dimensional nanoplatelets (NPLs) and nanocrystal cubes. The Stokes shift size dependence trend of NPLs and nanocubes is non-monotonic indicating different physics is at play for the two nanocrystal morphologies. For nanocrystals with cubic shape, with the increase of edge length, there is a significant decrease in Stokes shift values. However, for NPLs with the increase of thickness (1–4 ML), Stokes shift values will increase. The study emphasizes the transition from a spectroscopic point of view and relates the two Stokes shift trends to 2D and 0D exciton dimensionalities for the two morphologies. Our findings highlight the importance of CsPbBr 3 nanocrystal morphology for Stokes shift prediction.

36 MATERIALS SCIENCE↗