Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “big data applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Performance Analysis of Data Processing in Distributed File Systems with Near Data Processing

In the era of big data, the escalating volume and velocity of data generation pose significant challenges in data processing. Traditional systems like Spark and Hadoop manage the increasing amount and velocity of data by improving data placement and processing speeds. However, they face inherent limitations due to the essential data movement required for processing. In this paper, we explore the Skyhook framework, a novel extension of the Ceph distributed system, which significantly reduces the need for data movement. We present an extensive case study using the Skyhook framework, applying it with the TPC-H and K-means clustering algorithms. More specifically, we leverage the TPC-H benchmark to distinguish between CPU-intensive and I/O-intensive tasks. We explore the integration of K-means clustering into SQL, coupled with a near-data processing system to offload the computational burden of the K-means clustering algorithm to storage nodes. We conduct a comprehensive performance evaluation of distributed data processing applications across three processing approaches: traditional layout (baseline), optimized layout, and near-data processing. Additionally, we introduce the use of the FIO tool to simulate real-world system workloads, enabling the measurement of performance metrics such as average latency and CPU utilization. Our research is a significant advance in understanding how to optimize data processing systems to meet the demands of the modern data landscape.

Hou, Shiyue↗

Investigating resource-efficient neutron/gamma classification ML models targeting eFPGAs

There has been considerable interest and resulting progress in implementing machine learning (ML) models in hardware over the last several years from the particle and nuclear physics communities. A big driver has been the release of the Python package, hls4ml, which has enabled porting models specified and trained using Python ML libraries to register transfer level (RTL) code. So far, the primary end targets have been commercial field-programmable gate arrays (FPGAs) or synthesized custom blocks on application specific integrated circuits (ASICs). However, recent developments in open-source embedded FPGA (eFPGA) frameworks now provide an alternate, more flexible pathway for implementing ML models in hardware. These customized eFPGA fabrics can be integrated as part of an overall chip design. In general, the decision between a fully custom, eFPGA, or commercial FPGA ML implementation will depend on the details of the end-use application. In this work, we explored the parameter space for eFPGA implementations of fully-connected neural network (fcNN) and boosted decision tree (BDT) models using the task of neutron/gamma classification with a specific focus on resource efficiency. We used data collected using an AmBe sealed source incident on Stilbene, which was optically coupled to an OnSemi J-series silicon photomultiplier (SiPM) to generate training and test data for this study. We investigated relevant input features and the effects of bit-resolution and sampling rate as well as trade-offs in hyperparameters for both ML architectures while tracking total resource usage. The performance metric used to track model performance was the calculated neutron efficiency at a gamma leakage of 10 -3 . The results of the study will be used to aid the specification of an eFPGA fabric, which will be integrated as part of a test chip.

47 OTHER INSTRUMENTATION↗

Data Science and Computation for Rapid and Dynamic Compression Experiment Workflows at Experimental Facilities, September 8-11, 2020. Workshop Report

The application of high pressure to materials has enabled discoveries in scientific fields such as planetary science, materials science, and materials synthesis. Recent advances in X-ray user light sources and other facilities, co-location and integration of user facilities with high-pressure drivers, availability of high-performance computing (HPC) platforms, and the development of new data science techniques have created opportunities for, and challenges in, advancing data analytics for rapid and dynamic compression experiments. To address these challenges, harness the emerging technology now available, and expedite scientific discovery, Los Alamos National Laboratory (LANL) hosted a virtual workshop entitled “Data Science and Computation for Rapid and Dynamic Compression Workflows at Experimental Facilities” from September 8 to 11, 2020. The workshop included 95 registered scientists and analytics experts from 15 universities, 9 United States (US) national laboratories, 5 US and European X-ray light sources, neutron sources such as the Los Alamos Neutron Science Center (LANSCE), other big science facilities such as the National Ignition Facility (NIF), and an industry representative. The workshop included 31 invited talks and 4 lightning talks by students and postdocs.

36 MATERIALS SCIENCE↗

Predictive Data-driven Platform for Subsurface Energy Production

Subsurface energy activities such as unconventional resource recovery, enhanced geothermal energy systems, and geologic carbon storage require fast and reliable methods to account for complex, multiphysical processes in heterogeneous fractured and porous media. Although reservoir simulation is considered the industry standard for simulating these subsurface systems with injection and/or extraction operations, reservoir simulation requires spatio-temporal “Big Data” into the simulation model, which is typically a major challenge during model development and computational phase. In this work, we developed and applied various deep neural network-based approaches to (1) process multiscale image segmentation, (2) generate ensemble members of drainage networks, flow channels, and porous media using deep convolutional generative adversarial network, (3) construct multiple hybrid neural networks such as convolutional LSTM and convolutional neural network-LSTM to develop fast and accurate reduced order models for shale gas extraction, and (4) physics-informed neural network and deep Q-learning for flow and energy production. We hypothesized that physicsbased machine learning/deep learning can overcome the shortcomings of traditional machine learning methods where data-driven models have faltered beyond the data and physical conditions used for training and validation. We improved and developed novel approaches to demonstrate that physics-based ML can allow us to incorporate physical constraints (e.g., scientific domain knowledge) into ML framework. Outcomes of this project will be readily applicable for many energy and national security problems that are particularly defined by multiscale features and network systems.

58 GEOSCIENCES↗

Machine Learning-Enabled Image Classification for Automated Electron Microscopy

Abstract Traditionally, materials discovery has been driven more by evidence and intuition than by systematic design. However, the advent of “big data” and an exponential increase in computational power have reshaped the landscape. Today, we use simulations, artificial intelligence (AI), and machine learning (ML) to predict materials characteristics, which dramatically accelerates the discovery of novel materials. For instance, combinatorial megalibraries, where millions of distinct nanoparticles are created on a single chip, have spurred the need for automated characterization tools. This paper presents an ML model specifically developed to perform real-time binary classification of grayscale high-angle annular dark-field images of nanoparticles sourced from these megalibraries. Given the high costs associated with downstream processing errors, a primary requirement for our model was to minimize false positives while maintaining efficacy on unseen images. We elaborate on the computational challenges and our solutions, including managing memory constraints, optimizing training time, and utilizing Neural Architecture Search tools. The final model outperformed our expectations, achieving over 95% precision and a weighted F-score of more than 90% on our test data set. This paper discusses the development, challenges, and successful outcomes of this significant advancement in the application of AI and ML to materials discovery.

Materials Science↗

Understanding the Thermal Physics and Metallurgy of Metal Big Area Additive Manufacturing

The research goal of this EPSCoR-DOE partnership is to mitigate defects in parts made using a new type of additive manufacturing (AM) process called metal Big Area Additive Manufacturing (m-BAAM). To realize this goal, the PIs will detect and correct defects in the part as it is being printed by combining fundamental knowledge of the thermal physics and metallurgy of m-BAAM with in-process sensor data. Developed at the DOE-funded Manufacturing Demonstration Facility at Oak Ridge National Laboratory, the m-BAAM process involves one or more robots working together to produce a part by fusing metal wire layer-by-layer using arc welding. The process can print large metal parts such as turbine blades, which is not possible using other AM processes. In addition, m-BAAM production rates are more than ten times faster than other AM processes while requiring one-tenth of the material cost. Despite their potential to become a critical force multiplier in the energy generation industry, m-BAAM parts may fail to print accurately due to retention of heat and uneven cooling. Overheating and anomalous cooling rates in turn can cause inconsistencies in the microstructure, leading to sudden failure when used in safety-critical applications. In other words, flaw formation in m-BAAM parts is governed by the thermal history – intensity and spatial distribution of heat inside the part during printing. The thermal history is a complex function of the part shape and process settings such as welding energy, path taken by the welding torch for deposition (tool path), wire feed rate, among others.

36 MATERIALS SCIENCE↗

In situ feature analysis for large-scale multiphase flow simulations

The study of multiphase flow is essential for designing chemical reactors such as fluidized bed reactors (FBR), as a detailed understanding of hydrodynamics is critical for optimizing reactor performance and stability. An FBR allows scientists to conduct different types of chemical reactions involving multiphase materials, especially interaction between gas and solids. During such complex chemical processes, the formation of void regions in the reactor, generally termed as bubbles, is an important phenomenon. The study of these bubbles has a deep implication in predicting the reactor’s overall efficiency. But physical experiments needed to understand bubble dynamics are costly and non-trivial due to the technical difficulties involved and harsh working conditions of the reactors. Therefore, to study such chemical processes and bubble dynamics, a state-of-the-art computational simulation MFIX-Exa is being developed. Despite the proven accuracy of MFIX-Exa in modeling bubbling phenomena, the large-scale output data prohibits the use of traditional post hoc analysis capabilities in both storage and I/O time. Herein, to address these issues and allow the application scientists to explore the bubble dynamics in an efficient and timely manner, we have developed an end-to-end analytics pipeline that enables in situ detection of bubbles, followed by a flexible post hoc visual exploration methodology of bubble dynamics. The proposed method enables interactive analysis of bubbles, along with quantification of several bubble characteristics, enabling experts to understand the bubble interactions in detail. Positive feedback from the experts has indicated the efficacy of the proposed approach for exploring bubble dynamics in very-large-scale multiphase flow simulations.

97 MATHEMATICS AND COMPUTING↗

A Data Science and Machine Learning Platform Supporting Large Particle Accelerator Control and Diagnostics Applications Final Report: SBIR Initial Phase II DE-SC0022583

The Machine Learning Data Platform (MLDP) is a product providing full-stack support for data science, Machine Learning, and Artificial Intelligence (ML/AI) applications at particle accelerator and large experimental physics facilities. It supports ML/AI applications from front-end, high-speed acquisition of heterogeneous, time-series data, through data archiving and management, to back-end analysis. The MLDP embodies a “data-science ready” platform for data analysis and ML/AI applications in diagnosis, modelling, control, and optimization of these facilities. It provides data scientists and applications a consistent, datacentric interface to archive data standardizing implementation and deployment of ML/AI algorithms to different operations configurations within the same facility, or between facilities. Being an open-source, public-domain project, the MLDP is intended for broadest possible impact by increasing accessibility and minimizing the required expertise for installation and operation. The MLDP can also be deployed at user facilities for experimental data collection, archiving, and analysis. It is capable of acquisition and archiving of heterogeneous data from experimental equipment (e.g., images, arrays, structures, etc.) along with system hardware configurations (e.g., scalars, tables), control system process variables, and any metadata required for provenance. Thus, the MLDP can manage experimental data through its entire lifecycle, from acquisition and archiving, through analysis and investigation, to release and final publication.

43 PARTICLE ACCELERATORS↗

Using satellite data to assess Hurricane Helene’s impact on vegetation by land use in the CSRA

Many regions within Georgia, South Carolina, and North Carolina experienced record-breaking rainfall and catastrophic winds due to Hurricane Helene. Helene made landfall on Florida’s big bend on September 26 th , 2024, as a category 4 hurricane, and tracked northward through Georgia and the southern Appalachian Mountains before dissipating on September 29 th , 2024. One significant impact of the hurricane was severe damage to tree canopies across the southeastern United States. This study utilizes satellite-based remotely sensed data provided by the National Aeronautical and Space Administration (NASA) to examine the resilience of these tree canopies following the hurricane. Specifically, the Normalized Difference Vegetation Index (NDVI), Leaf Area Index (LAI), Land Cover Type, Soil Moisture Active Passive (SMAP), and the Global Precipitation Measurement (GPM) mission datasets were employed to study the tree canopy’s response to such events in the Central Savannah River Area (CSRA). A significant increase was observed in the LAI in November of 2024. This increase in LAI was likely influenced by warmer-than-average temperatures throughout October of 2024, along with a second record-breaking rainfall event in the CSRA on November 6th, 2024. The LAI increase was then broken down into land cover type to understand which areas contributed most. The complex nature and resilience of trees became evident, offering opportunities for further exploration and application in urban planning, emergency response, and environmental management.

54 ENVIRONMENTAL SCIENCES↗

CHARACTERIZATION ON ANISOTROPIC THERMAL CONDUCTIVITY FOR BIG AREA ADDITIVE MANUFACTURING WITH POLYMERS

Additive manufacturing with polymers has been used mainly for prototyping. A recent development of Big Area Additive Manufacturing (BAAM) at Oak Ridge National Laboratory has opened its applications in the mold and die industry. A numerical simulation and prediction for a mold heating performance requires accurate anisotropic thermal properties of the printed material, which are challenging to obtain, and often requires the use of multiple techniques. The transient plane source (TPS) technique has been widely used due to its ability to measure the thermal properties of an extensive range of materials (solids, liquids, powder). Despite the capability to characterize thermal conductivity k of isotropic and anisotropic materials, the measurements of latter materials are limited to the cases, where the samples have the same thermal conductivity k along x- and y-axis that form the radial plane. In this work, the method for a characterization of k in all three dimensions is developed, and the application of TPS is extended to the determination of thermal properties along the x-, y-, and z-axis individually. The materials are represented by additively manufactured polymers including polylactic acid (PLA) and styrene maleic anhydride (SMA). The developed method consists of (1) a determination of the heat capacity of the polymers by means of TPS in combination with the developed in this work data analysis procedure, (2) a machining three types of cylindrical samples from the same material, with the height corresponding either to x-, y-, or z-direction of printing, and (3) a determination of axial thermal conductivity employing anisotropic model and using previously determined heat capacity

Trofimov, Artem↗

Augmentation of WRF-Hydro to simulate overland-flow- and streamflow-generated debris flow susceptibility in burn scars

In steep wildfire-burned terrains, intense rainfall can produce large runoff that can trigger highly destructive debris flows. However, the ability to accurately characterize and forecast debris flow susceptibility in burned terrains using physics-based tools remains limited. Here, we augment the Weather Research and Forecasting Hydrological modeling system (WRF-Hydro) to simulate both overland and channelized flows and assess postfire debris flow susceptibility over a regional domain. We perform hindcast simulations using high-resolution weather-radar-derived precipitation and reanalysis data to drive non-burned baseline and burn scar sensitivity experiments. Our simulations focus on January 2021 when an atmospheric river triggered numerous debris flows within a wildfire burn scar in Big Sur – one of which destroyed California's famous Highway 1. Compared to the baseline, our burn scar simulation yields dramatic increases in total and peak discharge and shorter lags between rainfall onset and peak discharge, consistent with streamflow observations at nearby US Geological Survey (USGS) streamflow gage sites. For the 404 catchments located in the simulated burn scar area, median catchment-area-normalized peak discharge increases by ~ 450 % compared to the baseline. Catchments with anomalously high catchment-area-normalized peak discharge correspond well with post-event field-based and remotely sensed debris flow observations. We suggest that our regional postfire debris flow susceptibility analysis demonstrates WRF-Hydro as a compelling new physics-based tool whose utility could be further extended via coupling to sediment erosion and transport models and/or ensemble-based operational weather forecasts. Given the high-fidelity performance of our augmented version of WRF-Hydro, as well as its potential usage in probabilistic hazard forecasts, we argue for its continued development and application in postfire hydrologic and natural hazard assessments.

54 ENVIRONMENTAL SCIENCES↗

T-FSM: A Scalable Distributed Task-Based System for Frequent Subgraph Pattern Mining from a Big Graph

Finding frequent subgraph patterns in a big graph is an important problem with many applications such as classifying chemical compounds and building indexes to speed up graph queries. Since this problem is NP-hard, some recent parallel and distributed systems have been developed to accelerate the mining. However, they often have a huge memory cost, very long running time, suboptimal load balancing, poor scale-out capability, and possibly inaccurate results. In this article, we propose an efficient system called T-FSM for parallel mining of frequent subgraph patterns in a big graph. T-FSM supports a new anti-monotonic frequentness measure called Fraction-Score, which is more accurate than the widely used MNI measure. The execution engine of T-FSM supports both intra-machine parallelism and inter-machine parallelism. For intra-machine parallelism, T-FSM adopts a novel task-based execution model to ensure high multithreading concurrency, bounded memory consumption, and effective load balancing. For inter-machine parallelism, T-FSM ensures good scale-out performance with a lightweight pattern rebalancing approach that reduces workload skewness of pattern evaluations among machines. To avoid recomputing the contexts for migrated patterns, we design a novel context cache table to support concurrent and asynchronous requesting and caching of remote context data, which can timely evict and garbage collect used pattern contexts that are no longer needed to keep memory consumption bounded. Extensive experiments show that T-FSM is orders of magnitude faster than existing state-of-the-art parallel systems (more than 10×, 51×, 131×, 55× speedup over ScaleMine, DistGraph, Pangolin and Peregrine, respectively) and distributed systems (more than 42× and 88× over ScaleMine and DistGraph, respectively) for frequent subgraph pattern mining, and it scales out satisfactorily to 512 CPU cores on the Polaris supercomputer at Argonne National Laboratory.

97 MATHEMATICS AND COMPUTING↗

Tikiri—Towards a lightweight blockchain for IoT

Internet of Things (IoT) platforms have been deployed in several domains to enhance efficiency of business process and improve productivity. Most IoT platforms comprise of heterogeneous software and hardware components which can potentially introduce security and privacy challenges. Blockchain technology has been proposed as one of the solutions to realize IoT security by leveraging the (a) Immutable ledger, (b) Decentralized architecture and (c) Strong cryptography primitives. However, integrating blockchain platforms with IoT based applications presents several challenges due to lack of (a) acceptable performance on resource-constrained devices, (b) high transaction throughput, (c) keyword-based search and retrieve, (d) transaction back pressure operations, and (e) real-time response. In this paper, we propose a lightweight blockchain platform, “Tikiri”, for resource-constrained IoT devices. Tikiri uses Apache Kafka for the consensus and proposes new blockchain architecture to handle real-time transaction execution on the blockchain. Tikiri is characterized by functional programming and actor-based smart contract platform that realizes concurrent execution of transactions in the blockchain. Tikiri realizes a lightweight and scalable blockchain that can provides performance on the resource-constrained IoT devices.

97 MATHEMATICS AND COMPUTING↗

A method for assessing economic, environmental, and reliability tradeoffs of interregional transmission connecting ERCOT (the Texas grid) to the eastern and western grids

Reliable development of the power grid is an evolving concern for humanity due to extreme weather that frequently threatens power sector infrastructure. The state of Texas is a uniquely structured testbed for grid planners to study when looking for solutions to development, innovation, and overcoming such challenges. Because of its size and islanded structure, Texas is small enough to model, but big enough to matter. Texas is a global leader in energy production, energy consumption, and maintains an unusually diverse fuel mix. In addition, the state has experienced winter freezes, heat waves, wind storms, droughts and floods that have threatened power sector infrastructure or caused recent blackouts and calls for demand side conservation. One of the most devastating of these events was the North American winter storm, dubbed “Winter Storm Uri” by the Weather Channel, that froze the region in February 2021 and led to an extended power outage event that put the majority of Texan residents in darkness for days. While preparing to avoid such outage events in the future, various tools have been proposed to improve grid reliability, including energy efficiency, demand response, and distributed energy resources. An additional option would be to develop interregional transmission that connects the Texas grid to other national grids. To assess the merits of this idea, we developed a novel, universally-applicable and internationally-relevant framework to study how the Texas grid would evolve alongside access to various interregional ties. This method allows us to stress the synthetic grid structure and analyze how it would respond to the shock of a simulated winter storm event. Our method leverages open-source modeling tools, such as PowerGenome, pyGRETA, and GenX to synthesize unique zonal grid data, construct a consolidated network of model regions, and simulate different developmental pathways of capacity expansion and operational dispatch. We demonstrate our method with an analysis connecting the Electric Reliability Council of Texas (ERCOT), the grid that serves most of Texas, the Western Electricity Coordinating Council (WECC), the grid that serves the western half of the contiguous U.S., and the Eastern Interconnect, the grid that serves the eastern half of the contiguous U.S. Our results indicate that the cost-optimal capacity of interregional transmission connecting the ERCOT grid to other grids lies between 9–13 GW assuming baseline conditions. Building this amount of connecting capacity in one or multiple directions lowers the costs and emissions of development and operation by up to $16 billion and 257 million metric tonnes (MMT) respectively. Additionally, our results show that the interregional connections between ERCOT and other national grids reduce the amount of total load shed required through mild winter storm events. However, our results also show that there is a threshold of very extreme winter storm conditions, spanning multiple service areas, above which the connections exacerbate resource adequacy problems. Therefore, the results indicate that the connections need to be carefully planned alongside the rest of the grid infrastructure to avoid over-reliance on specific resources or technology options.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Dimensionally reduced machine learning model for predicting single component octanol–water partition coefficients

Abstract MF-LOGP, a new method for determining a single component octanol–water partition coefficients ( $$LogP$$ LogP ) is presented which uses molecular formula as the only input. Octanol–water partition coefficients are useful in many applications, ranging from environmental fate and drug delivery. Currently, partition coefficients are either experimentally measured or predicted as a function of structural fragments, topological descriptors, or thermodynamic properties known or calculated from precise molecular structures. The MF-LOGP method presented here differs from classical methods as it does not require any structural information and uses molecular formula as the sole model input. MF-LOGP is therefore useful for situations in which the structure is unknown or where the use of a low dimensional, easily automatable, and computationally inexpensive calculations is required. MF-LOGP is a random forest algorithm that is trained and tested on 15,377 data points, using 10 features derived from the molecular formula to make $$LogP$$ LogP predictions. Using an independent validation set of 2713 data points, MF-LOGP was found to have an average $$RMSE$$ RMSE = 0.77 ± 0.007, $$MAE$$ MAE = 0.52 ± 0.003, and $${R}^{2}$$ R 2 = 0.83 ± 0.003. This performance fell within the spectrum of performances reported in the published literature for conventional higher dimensional models ( $$RMSE$$ RMSE = 0.42–1.54, $$MAE$$ MAE = 0.09–1.07, and $${R}^{2}$$ R 2 = 0.32–0.95). Compared with existing models, MF-LOGP requires a maximum of ten features and no structural information, thereby providing a practical and yet predictive tool. The development of MF-LOGP provides the groundwork for development of more physical prediction models leveraging big data analytical methods or complex multicomponent mixtures. Graphical Abstract

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Generation and representation of synthetic smart meter data

Advanced energy algorithms running at big-data scale will be necessary to identify, realize, and verify energy savings to meet government and utility goals of building energy efficiency. Any algorithm must be well characterized and validated before it is trusted to run at these scales. Smart meter data from real buildings will ultimately be required for the development, testing, and validation of these energy algorithms and processes. However, for initial development and testing, smart meter data are difficult to work with due to privacy restrictions, noise from unknown sources, data accessibility, and other concerns which can complicate algorithm development and validation. This paper describes a new methodology to generate synthetic smart meter data of electricity use in buildings using detailed building energy modeling, which aims to capture the variability and stochastics of real energy use in buildings. The methodology can create datasets tailored to represent specific scenarios with known truth and controllable amounts of synthetic noise. Knowledge of ground truth also allows the development and validation of enhanced processes which leverage building metadata, such as building type or size (floor area), in addition to smart meter data. The methodology described in this paper includes the key influencing factors of real-world building energy use including weather data, occupant-driven loads, building operation and maintenance practices, and special events. Data formats to support workflows leveraging both synthetic meter data and associated metadata are proposed and discussed. Finally, example use cases of the synthetic meter data are described to illustrate potential applications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Differentiable, Learnable, Regionalized Process-Based Models With Multiphysical Outputs can Approach State-Of-The-Art Hydrologic Prediction Accuracy

Predictions of hydrologic variables across the entire water cycle have significant value for water resources management as well as downstream applications such as ecosystem and water quality modeling. Recently, purely data-driven deep learning models like long short-term memory (LSTM) showed seemingly insurmountable performance in modeling rainfall runoff and other geoscientific variables, yet they cannot predict untrained physical variables and remain challenging to interpret. Here, we show that differentiable, learnable, process-based models (called δ models here) can approach the performance level of LSTM for the intensively observed variable (streamflow) with regionalized parameterization. We use a simple hydrologic model HBV as the backbone and use embedded neural networks, which can only be trained in a differentiable programming framework, to parameterize, enhance, or replace the process-based model's modules. Without using an ensemble or post-processor, δ models can obtain a median Nash-Sutcliffe efficiency of 0.732 for 671 basins across the USA for the Daymet forcing data set, compared to 0.748 from a state-of-the-art LSTM model with the same setup. For another forcing data set, the difference is even smaller: 0.715 versus 0.722. Meanwhile, the resulting learnable process-based models can output a full set of untrained variables, for example, soil and groundwater storage, snowpack, evapotranspiration, and baseflow, and can later be constrained by their observations. Both simulated evapotranspiration and fraction of discharge from baseflow agreed decently with alternative estimates. The general framework can work with models with various process complexity and opens up the path for learning physics from big data.

54 ENVIRONMENTAL SCIENCES↗

Applications and Techniques for Fast Machine Learning in Science

In this community review report, we discuss applications and techniques for fast machine learning (ML) in science—the concept of integrating powerful ML methods into the real-time experimental data processing loop to accelerate scientific discovery. The material for the report builds on two workshops held by the Fast ML for Science community and covers three main areas: applications for fast ML across a number of scientific domains; techniques for training and implementing performant and resource-efficient ML algorithms; and computing architectures, platforms, and technologies for deploying these algorithms. We also present overlapping challenges across the multiple scientific domains where common solutions can be found. This community report is intended to give plenty of examples and inspiration for scientific discovery through integrated and accelerated ML solutions. This is followed by a high-level overview and organization of technical advances, including an abundance of pointers to source material, which can enable these breakthroughs.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗