Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Large Dataset Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Using Multi-scale Modeling System to Study the Interactions between Clouds, Precipitation, Aerosols, Radiation and Land Surface

Numerical cloud models, which are based the non-hydrostatic equations of motion, have been extensively applied to cloud-scale and mesoscale processes during the past four decades. Because cloud-scale dynamics are treated explicitly, uncertainties stemming from convection that have to be parameterized in (hydrostatic) large-scale models are obviated, or at least mitigated, in cloud models. Global models will use the non-hydrostatic framework when their horizontal resolution becomes about 10 kilometers, the theoretical limit for the hydrostatic approximation. This juncture will be reached one to two decades from now. Over the past generation, voluminous datasets on atmospheric convection have been accumulated from radar, instrumented aircraft, satellites, and rawinsonde measurements in field campaigns, enabling the detailed evaluation of models. Improved numerical methods have resulted in more accurate and efficient dynamical cores in models. Improvements have been made in the parameterizations of microphysical processes, radiation, boundary-layer effects, and turbulence; however, microphysical parameterizations remain a major source of uncertainty in all classes of atmospheric models. In recent years, exponentially increasing computer power has extended cloud-resolving-model integrations from hours to months, the number of computational grid points from less than a thousand to close to ten million. Three-dimensional models are now more prevalent. Much attention is devoted to precipitating cloud systems where the crucial-lkm scales are resolved in horizontal domains as large as 10,000 kilometers in two-dimensions, and 1,000 x 1,000 square kilometers in three-dimensions. Cloud models now provide statistical information useful for developing more realistic physically based parameterizations for climate models and numerical weather prediction models. It is also expected that NWP and mesoscale model can be run in grid size similar to cloud resolving model through nesting technique. A review of developments, improvements and applications of cloud models (GCE and WRF) at Goddard wlll be is presented in this talk. In particular, a new approach to using multi-scale modeling system to study the interactions between clouds, precipitation, aerosols and land will be presented.

Tao, Wei-Kuo↗

Using Multi-scale Modeling System to Study the Interactions between Clouds, Precipitation, Aerosols, Radiation and Land Surface

Numerical cloud models, which are based the non-hydrostatic equations of motion, have been extensively applied to cloud-scale and mesoscale processes during the past four decades. Because cloud-scale dynamics are treated explicitly, uncertainties stemming from convection that have to be parameterized in (hydrostatic) large-scale models are obviated, or at least mitigated, in cloud models. Global models will use the non-hydrostatic framework when their horizontal resolution becomes about 10 km, the theoretical limit for the hydrostatic approximation. This juncture will be reached one to two decades from now. Over the past generation, voluminous datasets on atmospheric convection have been accumulated from radar, instrumented aircraft, satellites, and rawinsonde measurements in field campaigns, enabling the detailed evaluation of models. Improved numerical methods have resulted in more accurate and efficient dynamical cores in models. Improvements have been made in the parameterizations of microphysical processes, radiation, boundary-layer effects, and turbulence; however, microphysical parameterizations remain a major source of uncertainty in all classes of atmospheric models. In recent years, exponentially increasing computer power has extended cloud-resolving-model integrations from hours to months, the number of computational grid points from less than a thousand to close to ten million. Three-dimensional models are now more prevalent. Much attention is devoted to precipitating cloud systems where the crucial 1-km scales are resolved in horizontal domains as large as 10,000 km in two-dimensions, and 1,000 x 1,000 sq km in three-dimensions. Cloud models now provide statistical information useful for developing more realistic physically based parameterizations for climate models and numerical weather prediction models. It is also expected that NWP and mesoscale model can be run in grid size similar to cloud resolving model through nesting technique. A review of developments, improvements and applications of cloud models (GCE and WRF) at Goddard will be presented in this talk. In particular, a new approach to using multi-scale modeling system to study the interactions between clouds, precipitation, aerosols and land will be presented.

Tao, Wei-Kuo↗

COMPASS-FME Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) Experiment Tree Inventory

This is the tree inventory (diameter, species, and live/dead status) data from the Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) experimental site. This manipulative, ecosystem-scale TEMPEST experiment is part of the COMPASS-FME (Coastal Observations, Mechanisms, and Predictions Across Systems and Scales: Field Measurements and Experiments; see https://compass.pnnl.gov/FME/COMPASSFME) project. It addresses the potential for freshwater and estuarine-water disturbance events to alter tree function, species composition, and ecosystem processes in a deciduous coastal forest in eastern Maryland, USA. The experiment uses a large-unit (2000 m2), un-replicated experimental design, with three 50 m × 40 m plots serving as control, freshwater, and estuarine-water treatments.This dataset includes:- An overall dataset README file.- The tree inventory data in both "wide" and "long" forms. These contain the same information but are structured differently, with the former more useful for human viewers and the latter more amenable for programmatic analyses.- A key to the species/genus codes used, which follow the U.S. Department of Agriculture's PLANTS schema (https://plants.usda.gov/).- A copy of the R code used to generate the wide- and long-form data files.All files are comma-separated value (CSV) and no special software is required to read them.

54 ENVIRONMENTAL SCIENCES↗

SymbolFit: Automatic Parametric Modeling with Symbolic Regression

We introduce SymbolFit (API: https://github.com/hftsoi/symbolfit), a framework that automates parametric modeling by using symbolic regression to perform a machine-search for functions that fit the data while simultaneously providing uncertainty estimates in a single run. Traditionally, constructing a parametric model to accurately describe binned data has been a manual and iterative process, requiring an adequate functional form to be determined before the fit can be performed. The main challenge arises when the appropriate functional forms cannot be derived from first principles, especially when there is no underlying true closed-form function for the distribution. In this work, we develop a framework that automates and streamlines the process by utilizing symbolic regression, a machine learning technique that explores a vast space of candidate functions without requiring a predefined functional form because the functional form itself is treated as a trainable parameter, making the process far more efficient and effortless than traditional regression methods. We demonstrate the framework in high-energy physics experiments at the CERN Large Hadron Collider (LHC) using five real proton-proton collision datasets from new physics searches, including background modeling in resonance searches for high-mass dijet, trijet, paired-dijet, diphoton, and dimuon events. We show that our framework can flexibly and efficiently generate a wide range of candidate functions that fit a nontrivial distribution well using a simple fit configuration that varies only by random seed, and that the same fit configuration, which defines a vast function space, can also be applied to distributions of different shapes, whereas achieving a comparable result with traditional methods would have required extensive manual effort.

Tsoi, Ho Fung [Univ. of Pennsylvania, Philadelphia↗

CMPLE: Correlation Modeling to Decode Photosynthesis Using the Minorize–Maximize Algorithm

In plant genomic experiments, correlations among various biological traits (phenotypes) give new insights into how genetic diversity may have tuned biological processes to enhance fitness under diverse conditions. Consequently, knowing how the correlations are affected by genetic (G) and environmental (E) factors helps develop climate-resilient plants. However, the current literature lacks any method for assessing the effect of predictors on pairwise correlations among multiple phenotypes together with easily interpretable model parameters. To address this need, we propose to model pairwise correlations directly in terms of G and E and develop a computationally efficient inference procedure. Two major novelties in our methodology are (1) the use of a composite pairwise likelihood method to avoid the positive definiteness restriction on the correlation matrix and (2) the use of a novel Minorize–Maximize (MM) algorithm for the efficient estimation of a large number of parameters. The proposed method shows excellent numerical performance on synthetic datasets. Here, the analysis of the motivating data on cowpea reveals that the rates of solar energy storage by photosynthesis (the aggregate trait) are differentially affected by different genetic loci through two distinct processes: “photoinhibition” which results from photodamage caused by excess light, and “photoprotection” which protects plants from photodamage but also results in energy loss.

Correlation modeling↗

The Artificial Scientist: in-Transit Machine Learning of Plasma Simulations

Large-scale simulations or scientific experiments produce petabytes of data per run. This poses massive challenges for I/O and storage when scientific analysis workflows are run manually offline. Unsupervised deep learning-based techniques to extract patterns and non-linear relations from these large amounts of data provide a way to build scientific understanding from raw data, reducing the need for manual pre-selection of analysis steps, but require exascale compute and memory to process the full dataset available. In this paper, we demonstrate a heterogeneous streaming workflow in which plasma simulation data is streamed directly to a Machine Learning (ML) application training a model on the simulation data in-transit, completely circumventing the capacity-constrained filesystem bottleneck. This workflow employs openPMD to provide a high level interface to describe scientific data and also uses ADIOS2, to transfer volumes of data that exceed the capabilities of the filesystem. We employ experience replay to avoid catastrophic forgetting in learning from this non-steady state process in a continual manner and adapt it to improve model convergence while learning in-transit. As a proof-of-concept, we approach the ill-posed inverse problem of predicting particle dynamics from radiation in a particle-incell (PIConGPU) simulation of the Kelvin-Helmholtz instability (KHI). We detail hardware-software co-design challenges as we scale PIConGPU to full Frontier, the Top-1 system as of June 2024 Top500 list.

Kelling, Jeffrey [Helmholtz-Zentrum Dresden Rossen↗

Airborne LiDAR to Improve Canopy Fuels Mapping for Wildfire Modeling

Increasing conflict between wildfire and the built environment has increased the need for more up-to-date and finer resolution canopy fuels data to improve wildfire modeling and associated risk forecasts. The US Forest Service and US Department of the Interior’s LANDFIRE product, which provides 30-m resolution canopy fuels data for the entire US, is one of the most widely used sources of fuels data. However, the last complete mapping effort for LANDFIRE is based on 2016 conditions, and subsequent updates reflect disturbances 1-2 years behind the release year. Airborne systems equipped with Light Detection and Ranging (LiDAR) sensors can be deployed to actively sense canopy structure and estimate canopy fuels data (cover, height, base height, bulk density) at finer resolutions. Canopy base height (CBH) and canopy bulk density (CBD) are difficult to measure both in the field and in LiDAR point clouds. Still, they are important for accurately modeling crown fires, which are often intense and difficult to contain. Additionally, point cloud datasets are large, and calculations require efficient utilization of computational resources. To address these challenges, we are working on an approach that uses openly available National Ecological Observatory Network (NEON) airborne LiDAR data, with calculations processed in the R programming language and parallelized through the lidR package. CBH and CBD are often derived from tree height, diameter at breast height, and species-specific allometries using the Fire and Fuels Extension of the Forest Vegetation Simulator (FFE-FVS). We aim to test if airborne LiDAR can estimate CBH and CBD without the use of empirical equations. Reliable estimates of canopy fuels data directly from airborne LiDAR could streamline quick, fine-resolution updates for use in wildfire behavior models.

54 ENVIRONMENTAL SCIENCES↗

Preliminary Design of Engineering-Scale Salt Accident Analysis Facility to Support Molten Salt Reactor Licensing

Systems-level nuclear accident analysis codes for reactor licensing must be validated using experimental data that represent behaviors expected during actual full-scale accidents. Some behaviors may arise from coupled processes and only manifest at large scales. This report presents the preliminary design of the Salt Accident Analysis Facility (SAAF, pronounced “safe”), which is an experimental test facility to be constructed at Argonne that can be used to conduct integrated salt accident tests at an engineering scale. The measurement capabilities of the SAAF are based on previously developed methods and will provide the representative datasets that are needed to support molten salt reactor (MSR) licensing. Details of the design, the processes to be quantified, the measurement techniques for quantifying the processes, the variables that can be adjusted to simulate different accident scenarios, and operational considerations are presented herein. This report provides stakeholders the opportunity to give feedback on the test facility capabilities and planned analyses before it is constructed.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Assessing the feasibility of a spaceborne 3D lightning observing concept

The distribution of electrical charge in thunderclouds results from thermodynamic, microphysical, and kinematic processes, which also modulate thunderstorm evolution. It is no surprise that the connection between lightning and these physical processes is so strong that the increase and vertical growth of lightning activity closely follows the vertical growth of the thundercloud, but unraveling these connections is not trivial and requires observations of the three-dimensional (3D) structure of electrical activity in a cloud. Ground-based 3D lightning mapping networks give excellent 3D flash-level detail but are limited to regional coverage. Satellite-based optical lightning mappers give excellent global coverage but are largely limited to 2D summaries of flash rate and radiant intensity, albeit new flash products and stereographic techniques are chipping away this limitation. New observing strategies are needed to expand and diversify the corpus of 3D lightning datasets and motivate studies that unravel connections lightning has with these key physical processes and the surrounding environment. This study examines the feasibility of using a distributed network of orbing satellites with VHF-based lightning detectors to obtain global maps of 3D lightning activity and assess efficacy of this approach for use in a new, small satellite mission concept called CubeSpark. CubeSpark combines new VHF and high-resolution, bispectral optical instruments on a constellation of low-Earth orbiting (LEO) satellites to globally map the 3D electrical structure of thunderstorms and study how it relates to thunderstorm evolution, extreme weather, nitrogen oxide production and distribution, upper atmospheric electrical phenomena, and how 3D flash observations can complement existing satellite-based lightning mappers and improve decision support tools. To locate lightning discharges, CubeSpark seeks to use the VHF time-of-arrival technique, similar to ground-based total lightning mapping networks. The vertical location accuracy of these satellite retrievals will be of poorer quality compared to a ground-based network, which has non-trivial implications for lightning flash reconstruction and lightning-based interpretations of deep convection. We adapt a Lightning Mapping Array (LMA) simulation framework to an orbiting network and use it to address feasibility of 3D lightning detection from space with particular attention to location accuracy of VHF detections of lightning in the vertical. These simulations inform a constellation design study that defines a realistic orbital configuration and depicts the global coverage for CubeSpark. Results indicate that a 3D location accuracy of <1-2 km for each dimension can be achieved across 300-500 km wide swaths, which suggests that CubeSpark can resolve the charge structure of thunderclouds from the tropics to the mid- and high- latitudes.

Lightning↗

Automation of Laser Plasma Focused Ion Beam Microscopy for Next-Gen Energy Materials

Automation can revolutionize the use of ultrafast laser ablation and plasma-focused ion beam (PFIB) techniques for high-throughput, reproducible cross-sectioning and various sample preparation in materials characterization. As these methods become essential for analyzing complex energy materials and next-generation devices, efficient, standardized workflows are needed to minimize variability and enhance precision. This work highlights our advancements in developing automated processes for sample preparation that integrates machine learning, workflow optimization, and large-scale data acquisition to improve efficiency and scalability in applications such as electrolyzers, photovoltaic cells, and microelectronics. To streamline cross-sectioning and lamella fabrication, we have implemented fully automated workflows that standardize laser ablation and PFIB milling sequences. These workflows incorporate pre-programmed protocols for material removal, alignment, and thinning, reducing user intervention and ensuring consistency across different sample types. Machine learning algorithms further enhance automation by predicting optimal milling strategies and adapting parameters based on material properties and sectioning requirements. This approach significantly improves throughput while maintaining the structural integrity of prepared samples for high-resolution imaging and analysis, including transmission electron microscopy. Beyond sample preparation, our automation platform enables the acquisition of large, high-resolution datasets through serial sectioning, image alignment, and 3D reconstruction. These automated routines facilitate multi-scale characterization, capturing structural and compositional details from the nanoscale to the device level. By reducing variability and increasing efficiency, our automated approach enhances defect analysis, failure diagnostics, and process optimization, accelerating advancements in materials research and device engineering.

36 MATERIALS SCIENCE↗

An AI-driven framework for evaluating local and state authorities’ permitting processes

The demand for new energy infrastructure is increasing across the United States, but heterogenous permitting processes and embedded requirements across different local jurisdictions can cause project delays, increase “soft costs,” and hinder developer expansion. This study analyzes the variability in local permitting requirements across the U.S. and develops a quantitative approach to describe their clarity and effectiveness in enabling infrastructure project development. By using an Energy Language Model (ELM), a large language model (LLM) for energy technologies, we systematically gathered permitting information from nearly 300 state-, county-, and city-level documents, creating a structured dataset of requirements and procedures on an unprecedented scale and speed. Our analysis revealed that local (city and county) permitting requirement documents are underrepresented compared to state-level guidance documents, which can impede timely and cost-effective installation of new electric infrastructure. Our validation process showed that the final database has an accuracy of approximately 95%. We, further, created a new quantitative method to score permitting requirements for clarity and efficiency, with electric vehicle supply equipment as an initial use case. The average local permitting document scored a 1.8 out of 5, which we interpret as meaning that half of the requirements developers face when installing electric infrastructure are ambiguous, increasing both cost and time. We also created a “Generalized Permit Process”, highlighting common procedural steps and identifying specific opportunities for municipalities to improve their documentation. This research establishes a systematic and scalable framework for evaluating the complexities of local infrastructure permitting processes by combining LLM-powered data collection and quantitative scoring. The framework enables policymakers and developers to identify and mitigate procedural bottlenecks, with the expectation that these improvements can accelerate application review and approval, reduce project costs, and expedite connection to utility distribution grids. As a foundational approach for streamlining local project development processes, this study’s methods are intended to be extended to a wide range of energy applications.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Unsupervised multimodal fusion of in-process sensor data for advanced manufacturing process monitoring

Effective monitoring of manufacturing processes is crucial for maintaining product quality and operational efficiency. Modern manufacturing environments often generate vast amounts of complementary multimodal data, including visual imagery from various perspectives and resolutions, hyperspectral data, and machine health monitoring information such as actuator positions, accelerometer readings, and temperature measurements. However, fusing and interpreting this complex, high-dimensional data presents significant challenges, particularly when labeled datasets are unavailable or impractical to obtain. This paper presents a novel approach to multimodal sensor data fusion in manufacturing processes, inspired by the Contrastive Language-Image Pre-training (CLIP) model. We leverage contrastive learning techniques to correlate different data modalities without the need for labeled data, overcoming limitations of traditional supervised machine learning methods in manufacturing contexts. Our proposed method demonstrates the ability to handle and learn encoders for five distinct modalities: visual imagery, audio signals, laser position (x and y coordinates), and laser power measurements. By compressing these high-dimensional datasets into low-dimensional representational spaces, our approach facilitates downstream tasks such as process control, anomaly detection, and quality assurance. The unsupervised nature of our method makes it broadly applicable across various manufacturing domains, where large volumes of unlabeled sensor data are common. We evaluate the effectiveness of our approach through a series of experiments, demonstrating its potential to enhance process monitoring capabilities in advanced manufacturing systems. This research contributes to the field of smart manufacturing by providing a flexible, scalable framework for multimodal data fusion that can adapt to diverse manufacturing environments and sensor configurations. The proposed method paves the way for more robust, data-driven decision-making in complex manufacturing processes.

Contrastive Learning↗

Phosphorus sorption and its environmental predictors across pantropical forest soils sampled over the past decade

Tropical forest productivity is frequently constrained by soil phosphorus (P) availability, yet global Land Surface Model (LSM), which are used to simulate ecosystem processes, still represent P cycling in tropical regions only in a limited way, largely because of scarce observational data. Phosphorus adsorption and desorption of dissolved inorganic P to and from soil minerals (hereafter termed sorption), is an important process for predicting how much P is available to plants. This dataset was created to improve predictions of soil P sorption in tropical soils by identifying the isotherm equation that best describes pantropical soils. It includes raw measurements of environmental variables, such as soil properties and climate, together with P sorption data collected from 40 forest soil pits from 9 Forest Global Earth Observatory (ForestGEO) sites across 7 tropical countries during 2018-2022. The data are organized by site, country, and continent. Each site may include several soil pits. For each pit, P sorption was measured across a range of soil P concentrations to build sorption isotherm curves, typically with about 6 to 8 measurements per curve. While sorbed P varies across these concentration levels, the other environmental variables remain constant at the plot level.

Aluminum oxide↗

Improving Sim-to-Real Transfer in Vision-Based Robot Navigation Via Instance-Level GAN-Based Data Augmentation

Achieving robust vision-based robotic tasks requires large amounts of data, which are often difficult to obtain in real-world scenarios. Simulators and synthetic data offer a cost-effective alternative, but the visual gap between simulation and reality hinders the performance of models when deployed in real-world environments. In this paper, we present a data augmentation pipeline that integrates a foundation model (Segment Anything Model) with an unsupervised image-to-image translation model (CycleGAN) for instance-level domain transfer from simulation to reality. This pipeline enables the generation of realistic labeled data from synthetic images for training supervised machine learning models in vision-based navigation tasks. We evaluate our approach on real-world data for ego-vehicle pose estimation, a critical autonomous navigation task involving the prediction of cross-track position and heading angle relative to road center line markings. The results of our tests show that our GAN-based data augmentation pipeline significantly outperforms models trained solely on simulation data or on data processed with standard image augmentation methods for sim-to-real transfer, enhancing model robustness and generalizability in real-world scenarios. Our method provides a scalable and flexible data augmentation tool for leveraging large synthetic datasets to enhance vision-based robotic navigation tasks.

artificial intelligence↗

Challenges in Remote-Sensing of Hail: Examining the Performance and Biases of Satellite Hail Retrievals Using Aqua MODIS Visible/IR and AMSR-E Passive-Microwave Observations

Hail poses threats to myriad aspects of human life and society, infrastructure, and agriculture. Scientifically, hail can often cause large errors in precipitation retrieval and estimation, posing challenges to establishing the current climatology of severe storms and their future trend in a changing Earth system. Fortunately, hailstorms exhibit distinct signatures in spaceborne remote-sensing datasets (e.g. overshooting cloud tops in visible/IR, or brightness temperature depressions in passive-microwave imagery). Approaches that leverage these signatures, however, are not without their pitfalls,: passive-microwave channels have large footprints and exhibit non-uniform beam filling. Visible/IR instruments have fine horizontal resolution but are limited by their insensitivity to processes occurring below cloud top. Large horizontal areas of smaller scatterers may also meaningfully lower the brightness temperatures, especially if they are able to occupy large portions of the footprint. Radiative transfer simulations show that low frequencies such as 19- and 37-GHz can be scattered to extremely low brightness temperatures by high concentrations of smaller (graupel-sized) ice scatterers, especially in larger features that are more likely to occupy the footprint, which may cause climatologies to overestimate the frequency severe hail. To address this, we investigate the nearly simultaneous and colocated MODIS (visible/IR) and AMSR-E (passive-microwave) onboard the Aqua satellite to leverage both datasets together, pairing AMSR-E and MODIS signatures of severe convection with ground-based weather radar, severe weather reports, and environmental parameters defined by the MERRA-2 reanalysis over CONUS, and then explore the performance and challenges of the algorithm when we expand outside the United States into six different geographical regimes throughout the Aqua domain.

Sarah D Bang↗

Harnessing large language models’ zero-shot and few-shot learning capabilities for regulatory research

Abstract Large language models (LLMs) are sophisticated AI-driven models trained on vast sources of natural language data. They are adept at generating responses that closely mimic human conversational patterns. One of the most notable examples is OpenAI's ChatGPT, which has been extensively used across diverse sectors. Despite their flexibility, a significant challenge arises as most users must transmit their data to the servers of companies operating these models. Utilizing ChatGPT or similar models online may inadvertently expose sensitive information to the risk of data breaches. Therefore, implementing LLMs that are open source and smaller in scale within a secure local network becomes a crucial step for organizations where ensuring data privacy and protection has the highest priority, such as regulatory agencies. As a feasibility evaluation, we implemented a series of open-source LLMs within a regulatory agency’s local network and assessed their performance on specific tasks involving extracting relevant clinical pharmacology information from regulatory drug labels. Our research shows that some models work well in the context of few- or zero-shot learning, achieving performance comparable, or even better than, neural network models that needed thousands of training samples. One of the models was selected to address a real-world issue of finding intrinsic factors that affect drugs' clinical exposure without any training or fine-tuning. In a dataset of over 700 000 sentences, the model showed a 78.5% accuracy rate. Our work pointed to the possibility of implementing open-source LLMs within a secure local network and using these models to perform various natural language processing tasks when large numbers of training examples are unavailable.

Biochemistry & Molecular Biology↗

Spectroscopy-guided discovery of three-dimensional structures of disordered materials with diffusion models

Spectroscopy techniques such as x-ray absorption near edge structure (XANES) provide valuable insights into the atomic structures of materials, yet the inverse prediction of precise structures from spectroscopic data remains a formidable challenge. In this study, we introduce a framework that combines generative artificial intelligence models with XANES spectroscopy to predict three-dimensional atomic structures of disordered systems, using amorphous carbon (a-C) as a model system. In this work, we introduce a new framework based on the diffusion model, a recent generative machine learning method, to predict 3D structures of disordered materials from a target property. For demonstration, we apply the model to identify the atomic structures of a-C as a representative material system from the target XANES spectra. We show that conditional generation guided by XANES spectra reproduces key features of the target structures. Furthermore, we show that our model can steer the generative process to tailor atomic arrangements for a specific XANES spectrum. Finally, our generative model exhibits a remarkable scale-agnostic property, thereby enabling generation of realistic, large-scale structures through learning from a small-scale dataset (i.e. with small unit cells). Our work represents a significant stride in bridging the gap between materials characterization and atomic structure determination; in addition, it can be leveraged for materials discovery in exploring various material properties as targeted.

36 MATERIALS SCIENCE↗

LCLS Big Data Handling – How I Learned to Stop Worrying and Love the Data Deluge

Advanced data and computing systems are vital to Linac Coherent Light Source (LCLS) operations, data interpretation and overall scientific productivity. The transition to MHz-era operation marks a fundamental change in scale that requires new infrastructure and architectures to link LCLS to the required scale of computing needed for scientific interpretation. The LCLS-II Data System meets big data challenges by implementing configurable data reduction that can adapt to multiple science areas, real-time analysis frameworks to provide visualization and fast feedback, and the ability to transfer data to local and remote computational facilities for near real time analysis at the appropriate scale. Feature extracted information generated in the data analysis pipeline - at the edge, local compute, or remote High-Performance Computing (HPC) resources - can be used to steer experiments and inform user decisions during beam time. Artificial Intelligence and Machine Learning (AI/ML) techniques present new opportunities to rapidly analyse large datasets and direct experiments, but create new challenges in scaling, adaptability, complexity, and trustworthiness. We describe how the LCLS-II Data System architecture addresses its data-driven challenges in the areas of data acquisition, data processing, data management, and workflow orchestration to decrease the overall time-to-science and provide a vision for future developments.

artificial intelligence↗