Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “task allocation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Autonomous Fueling System for Heavy-Duty Fuel Cell Electric Trucks

The motivation for this project stemmed from the challenges associated with rapidly refueling heavy-duty hydrogen fuel cell electric trucks (FCETs). Current manual refueling processes for fast refueling involve large, heavy equipment (e.g., hoses three times heavier than standard) and pose ergonomic risks and potential for equipment damage. The goal was to develop and test an autonomous fueling system to improve ergonomics, enhance safety, increase equipment durability through design improvements, and potentially speed up the fueling process. This project aimed to add to the understanding of autonomous systems in the context of heavy-duty hydrogen refueling, evaluating the technical effectiveness of potential concepts. A successfully developed system would benefit the public by facilitating the adoption of zero-emission heavy-duty transport, reducing reliance on manual labor for a physically demanding task, and potentially improving the safety and efficiency of hydrogen refueling infrastructure. The major accomplishment during the project's active period was the completion of the system-level architecture task. This involved establishing a detailed list of system requirements covering interfaces, environmental conditions, regulatory compliance, industry standards, safety, security, performance capabilities, and optional features. Five key use cases for the autonomous system were also identified. However, due to internal restructuring at Nikola, the necessary resources could not be allocated to continue the project. Consequently, Nikola opted to discontinue the project. The award was mutually terminated by Nikola and the DOE.

08 HYDROGEN↗

Enabling Scalable and Extensible Memory-mapped Datastores in Userspace

Exascale workloads are expected to incorporate data-intensive processing in close coordination with traditional physics simulations. These emerging scientific, data-analytics and machine learning applications need to access a wide variety of datastores in flat files and structured databases. Programmer productivity is greatly enhanced by mapping datastores into the application process's virtual memory space to provide a unified “in-memory” interface. Currently, memory mapping is provided by system software primarily designed for generality and reliability. However, scalability at high concurrency is a formidable challenge on exascale systems. Also, there is a need for extensibility to support new datastores potentially requiring HPC data transfer services. In this article, we present UMap , a scalable and extensible userspace service for memory-mapping datastores. Furthermore, through decoupled queue management, concurrency aware adaptation, and dynamic load balancing, UMap enables application performance to scale even at high concurrency. We evaluate UMap in data-intensive applications, including sorting, graph traversal, database operations, and metagenomic analytics. Our results show that UMap as a userspace service outperforms an optimized kernel-based service across a wide range of intra-node concurrency by 1.22-1.9 × . We performed two case studies to demonstrate UMap 's extensibility. First, a new datastore residing in remote memory is incorporated into UMap as an application-specific plugin. Second, we present a persistent memory allocator Metall built atop UMap for unified storage/memory.

97 MATHEMATICS AND COMPUTING↗

The integration of heterogeneous resources in the CMS Submission Infrastructure for the LHC Run 3 and beyond

While the computing landscape supporting LHC experiments is currently dominated by x86 processors at WLCG sites, this configuration will evolve in the coming years. LHC collaborations will be increasingly employing HPC and Cloud facilities to process the vast amounts of data expected during the LHC Run 3 and the future HL-LHC phase. These facilities often feature diverse compute resources, including alternative CPU architectures like ARM and IBM Power, as well as a variety of GPU specifications. Using these heterogeneous resources efficiently is thus essential for the LHC collaborations reaching their future scientific goals. The Submission Infrastructure (SI) is a central element in CMS Computing, enabling resource acquisition and exploitation by CMS data processing, simulation and analysis tasks. The SI must therefore be adapted to ensure access and optimal utilization of this heterogeneous compute capacity. Some steps in this evolution have been already taken, as CMS is currently using opportunistically a small pool of GPU slots provided mainly at the CMS WLCG sites. Additionally, Power9 processors have been validated for CMS production at the Marconi-100 cluster at CINECA. This note will describe the updated capabilities of the SI to continue ensuring the efficient allocation and use of computing resources by CMS, despite their increasing diversity. The next steps towards a full integration and support of heterogeneous resources according to CMS needs will also be reported.

Pérez-Calero Yzquierdo, Antonio↗

FLEET: Flexible Efficient Ensemble Training for Heterogeneous Deep Neural Networks

Parallel training of an ensemble of Deep Neural Networks (DNN) on a cluster of nodes is an effective approach to shorten the process of neural network architecture search and hyper-parameter tuning for a given learning task. Prior efforts have shown that data sharing, where the common preprocessing operation is shared across the DNN training pipelines, saves computational resources and improves pipeline efficiency. Data sharing strategy, however, performs poorly for a heterogeneous set of DNNs where each DNN has varying computational needs and thus different training rate and convergence speed. This paper proposes FLEET, a flexible ensemble DNN training framework for efficiently training a heterogeneous set of DNNs. We build FLEET via several technical innovations. We theoretically prove that an optimal resource allocation is NP-hard and propose a greedy algorithm to efficiently allocate resources for training each DNN with data sharing. We integrate data-parallel DNN training into ensemble training to mitigate the differences in training rates and introduce checkpointing into this context to address the issue of different convergence speeds. Experiments show that FLEET significantly improves the training efficiency of DNN ensembles without compromising the quality of the result.

Guan, Hui↗

Rich Dynamics of a General Producer–Grazer Interaction Model under Shared Multiple Resource Limitations

Organism growth is often determined by multiple resources interdependently. However, growth models based on the Droop cell quota framework have historically been built using threshold formulations, which means they intrinsically involve single-resource limitations. In addition, it is a daunting task to study the global dynamics of these models mathematically, since they employ minimum functions that are non-smooth (not differentiable). To provide an approach to encompass interactions of multiple resources, we propose a multiple-resource limitation growth function based on the Droop cell quota concept and incorporate it into an existing producer–grazer model. The formulation of the producer’s growth rate is based on cell growth process time-tracking, while the grazer’s growth rate is constructed based on optimal limiting nutrient allocation in cell transcription and translation phases. We show that the proposed model captures a wide range of experimental observations, such as the paradox of enrichment, the paradox of energy enrichment, and the paradox of nutrient enrichment. Together, our proposed formulation and the existing threshold formulation provide bounds on the expected growth of an organism. Moreover, the proposed model is mathematically more tractable, since it does not use the minimum functions as in other stoichiometric models.

59 BASIC BIOLOGICAL SCIENCES↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

Optimizing High-Throughput Inference on Graph Neural Networks at Shared Computing Facilities with the NVIDIA Triton Inference Server

Abstract With machine learning applications now spanning a variety of computational tasks, multi-user shared computing facilities are devoting a rapidly increasing proportion of their resources to such algorithms. Graph neural networks (GNNs), for example, have provided astounding improvements in extracting complex signatures from data and are now widely used in a variety of applications, such as particle jet classification in high energy physics (HEP). However, GNNs also come with an enormous computational penalty that requires the use of GPUs to maintain reasonable throughput. At shared computing facilities, such as those used by physicists at Fermi National Accelerator Laboratory (Fermilab), methodical resource allocation and high throughput at the many-user scale are key to ensuring that resources are being used as efficiently as possible. These facilities, however, primarily provide CPU-only nodes, which proves detrimental to time-to-insight and computational throughput for workflows that include machine learning inference. In this work, we describe how a shared computing facility can use the NVIDIA Triton Inference Server to optimize its resource allocation and computing structure, recovering high throughput while scaling out to multiple users by massively parallelizing their machine learning inference. To demonstrate the effectiveness of this system in a realistic multi-user environment, we use the Fermilab Elastic Analysis Facility augmented with the Triton Inference Server to provide scalable and high-throughput access to a HEP-specific GNN and report on the outcome.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Artificial Intelligence for (AI) Nuclear Security: Expert Perspectives on AI Priorities for the Office of International Nuclear Security

Artificial intelligence (AI) has the potential to transform nuclear security operations, offering opportunities to enhance effectiveness while simultaneously introducing new challenges. As AI technologies rapidly evolve, agencies across the United States Government (USG) are researching, implementing, and evaluating various AI models and systems. Given the broad capabilities and applications of these technologies, it is essential for each agency to identify and articulate those areas where it can make meaningful contributions aligned with its mission and expertise. To address this need for strategic focus, in late Fiscal Year 2025 (FY2025), the Office of International Nuclear Security (INS) established an AI Task Force (AITF) to gather input from subject matter experts (SMEs) regarding the most appropriate role INS could serve in researching, evaluating, or implementing AI for nuclear security. The AITF engaged 15 experts from national laboratories with backgrounds in cyber security, physical security, transport security, insider threat mitigation, nuclear engineering, human-systems engineering, and AI/ML development. This white paper summarizes the insights gathered from these SMEs and presents a potential roadmap for INS engagement with AI technologies. The recommendations outlined here are intended to inform INS leadership as they make strategic decisions about resource allocation and program direction in this rapidly evolving technological domain.

97 MATHEMATICS AND COMPUTING↗

matsim-agents v1.0

matsim-agents is a multi-agent AI framework for atomistic materials simulation and discovery. It orchestrates large language models (LLMs), machine-learned interatomic potentials (MLIPs), and DFT codes into a single agentic loop running on laptops and DOE leadership-class supercomputers. MULTI-AGENT ORCHESTRATION A LangGraph state machine with three nodes: a Planner that converts a natural-language research objective into structured tasks; an Executor that dispatches atomistic tools and loops until the queue is empty; and an Analyst that summarizes results into a human-readable report. State is checkpointed after every step and human-in-the-loop gates can be inserted at any edge. HYPOTHESIS-DRIVEN DISCOVERY CHAT An interactive REPL (matsim-agents chat) that couples LLM dialogue with atomistic simulation. Chemical formulas are automatically detected in conversation turns and trigger a full crystal-phase exploration: structure generation → relaxation → stability scoring → result injection back into the conversation, creating a closed hypothesis-refinement loop. CRYSTAL PHASE ENUMERATION Given a composition, the phase explorer enumerates prototypes by stoichiometry: elemental (fcc/bcc/hcp/sc/diamond), binary 1:1 (rocksalt/CsCl/zincblende/ wurtzite/fluorite/rutile), ternary 1:1:3 (cubic perovskite), ternary 1:2:4 (perovskite + spinel), quaternary 1:1:2:6 (Fm-3m double perovskite). 2-D prototypes (graphene, h-BN, MoS2 2H/1T) and multilayer stacking are also supported via --include-2d and --num-layers. SUPERCELL GENERATION AND SITE DECORATION Auto-tiling to a minimum atom count (--min-atoms), explicit NxNxN tiling (--supercell), symmetry-distinct site decorations (--n-orderings), and isotropic lattice-scale sweeps (--lattice-scales) for volume bracketing. MLFF RELAXATION AND STABILITY SCORING HydraGNN (multi-headed GNN) drives structure relaxation via ASE with FIRE, BFGS, or BFGSLineSearch. Stability output: delta-E/atom ranking across phases and a max-residual-force dynamical-stability proxy. Other MLIPs (MACE, NequIP, Orb) can be plugged in through the same interface. DFT BACKENDS Quantum ESPRESSO pw.x and VASP 6.6 are first-class labellers. Both have validated GPU builds and SLURM/PBS launchers for three DOE platforms: Frontier (AMD MI250X, ROCm), Aurora (Intel PVC, oneAPI), Perlmutter (NVIDIA A100, CUDA). QE produces ~100 binaries (pw.x, ph.x, epw.x, ...). VASP supports scf, relax, vc-relax, and vc-relax-shape run types. ACTIVE-LEARNING LOOP matsim-agents al run CONFIG.yaml drives an iterative HydraGNN-DFT loop: MD generates candidates → ensemble/MC-dropout uncertainty selects the most informative → DFT labels them in parallel inside one allocation → dataset grows → HydraGNN retrains → repeat. DFT backend is a single YAML toggle (dft.backend: vasp | qe). LLM-generated seed structures are supported (no curated POSCAR library needed). Config uses ${VAR}, ${VAR:-default}, ${VAR:?msg} shell-style substitution for cross-user/cross-site portability. LLM BACKENDS Ollama (local, default), vLLM (HPC multi-GPU serving), OpenAI, Anthropic, HuggingFace Transformers+Accelerate. Selected at runtime via flag or env var with no code changes. HPC PORTABILITY Same Python entry points run on Frontier (ROCm 7.2), Aurora (oneAPI), and Perlmutter (CUDA 12). DFT and ML stacks are never co-loaded in the same shell; they couple through the scheduler and filesystem. Advanced multi-node launchers (serve, discovery-chat, single-relaxation, active-learning, QE warm-start) are provided for all three platforms. CODABENCH COMPETITION BUNDLE A self-contained benchmark: 159 atomistic test structures across 11 material classes, 5 tasks (formation energy, forces, ML relaxation, AI-DFT relaxation, phase stability ranking), public/private leaderboard split (30/70), and four ready-to-run baselines: MACE-MP-0, HydraGNN, UMA, AllScAIP.

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

Elastic Resource Management for Deep Learning Applications in a Container Cluster

The increasing demand for learning from massive datasets is restructuring our economy. Effective learning, however, involves nontrivial computing resources. Most businesses utilize commercial infrastructure providers (e.g., AWS) to host their computing clusters in the cloud, where various jobs compete for available resources. While cloud resource management is a fruitful research field that has made many advances in production, such as Kubernetes and YARN, few efforts have been invested to further optimize the system performance, especially for deep learning (DL) training jobs in a container cluster. This work introduces FlowCon, a system that is able to monitor the individual evaluation functions of DL jobs at runtime, and thus to make placement decisions on resource allocations elastically. Here, we present a detailed design and implementation of FlowCon and conduct intensive experiments over various DL models. The results demonstrate that FlowCon significantly improves DL job completion time and resource utilization efficiency, compared to default systems. According to the results, FlowCon is able to improve the completion time by up to 68.8% and meanwhile, reduce the makespan by 18.0%, in the presence of various DL job workloads.

97 MATHEMATICS AND COMPUTING↗

A lightweight method for evaluating in situ workflow efficiency

Performance evaluation is crucial to understanding the behavior of scientific workflows. In this study, we target an emerging type of workflow, called in situ workflows. These workflows tightly couple components such as simulation and analysis to improve overall workflow performance. To understand the tradeoffs of various configurable parameters for coupling these heterogeneous tasks, namely simulation stride, and component placement, separately monitoring each component is insufficient to gain insights into the entire workflow behavior. Through an analysis of the state-of-the-art research, we propose a lightweight metric, derived from a defined in situ step, for assessing resource usage efficiency of an in situ workflow execution. By applying this metric to a synthetic workflow, which is parameterized to emulate behaviors of a molecular dynamics simulation, we explore two possible scenarios (Idle Simulation and Idle Analyzer) for the characterization of in situ workflow execution. In addition to preliminary results from a recently published study [11], we further exploit the proposed metric to evaluate a practical in situ workflow with a real molecular dynamics application, i.e., GROMACS. Here, experimental results show that the in transit placement (analytics on dedicated nodes) sustains a higher frequency for performing in situ analysis compared to the helper-core configuration (analytics co-allocated with simulation).

97 MATHEMATICS AND COMPUTING↗

Robust Carbon Dioxide Plume Imaging Using Joint Tomographic Inversion of Seismic Onset Time and Distributed Pressure and Temperature Measurements (Final Report)

We develop and demonstrate rapid and cost-effective methodologies for spatiotemporal tracking of CO2 plumes during geologic sequestration using joint inversion of seismic data and distributed pressure and temperature measurements. Key elements of our methodology are: (a) a computationally efficient approach to pressure and temperature propagation, (b) analysis of time lapse seismic data using a novel ‘seismic onset time’ approach to detect fluid front propagation, and (c) data assimilation and uncertainty assessment via joint inversion of pressure, temperature and time lapse seismic data, and (d) validating the numerical tomographic inversion using a CO2 injection demonstration projects, specifically data collected from the from the Petra Nova Parish Holdings CCUS project in the West Ranch Field, Texas and the Chester-16 reef CO2 injection site in Northern Michigan which is part of the DOE Midwestern Carbon Sequestration Project. The research team is led by Texas A&M University and includes Battelle as a subcontractor with support from Shell, Anadarko, Chevron and JX Nippon. A carbon dioxide (CO2) water-alternating-gas (WAG) pilot was conducted to gain insights into tertiary oil recovery potential via CO2 flood in the West Ranch Field as part of the Petra Nova project, the world’s largest post-combustion CO2 capture and utilization initiative. With a fluvial formation geology and large contrasts in permeability, this is a challenging and novel application of CO2 enhanced oil recovery (EOR). We build a predictive dynamic model of the subsurface that incorporates the multiphase and compositional data acquired during the pilot operation. The calibrated model is used for the carbon dioxide plume imaging. The study began with an initialization of the pilot sector model extracted from a calibrated full-field model. The pilot model calibration follows a two-step hierarchical workflow. First, we performed a large-scale update of the permeability distribution by integrating available bottomhole pressure and multiphase production data. In the second step, local permeability field is fine-tuned using a streamline-based method to match CO2 breakthrough times at the producers. The predictive capability of the calibrated model was verified through two blind validation tests: (1) the model showed good agreement with saturation logs acquired at two observation wells; and (2) the model reproduced the CO2 recovery as a fraction of the injected CO2. The use of seismic onset times has shown great promise for integrating near-continuous seismic surveys for updating geologic models. In this study, we analyze the impact of seismic survey frequency on the onset time approach aiming to extend the application of onset time to infrequent seismic surveys. In addition, we quantitatively examine the nonlinearity of the onset time method and compare it to the commonly used amplitude inversion method. We carry out a sensitivity analysis of seismic survey frequency based on the complete seismic survey data (over 175 surveys) of steam injection in a heavy oil reservoir (Peace River Unit) in Canada. Our results show that an adequate onset time map can be obtained from the infrequent seismic surveys by interpolation between seismic surveys as long as there is no change in the dominant underlying physics between the successive surveys. The study also shows that nonlinearity of the onset time method can be -smaller than that of the amplitude inversion method by several orders of magnitude. Application to the Brugge benchmark case shows that the onset time method obtains comparable permeability update as the traditional seismic amplitude inversion method with faster computation and improved convergence characteristics. We extend the streamline-based data integration approach to incorporate distributed temperature sensor (DTS) data using the concept of thermal tracer travel time. Then, a hierarchical workflow composed of evolutionary and streamline methods is employed to jointly history match the DTS and pressure data. Finally, CO2 saturation and streamline maps are used to visualize the CO2 plume movement during the sequestration process. The hierarchical workflow is applied to a carbon sequestration project in a carbonate reef reservoir within the Northern Niagaran Pinnacle Reef Trend in Michigan, USA. The monitoring data set consists of distributed temperature sensing (DTS) data acquired at the injection well and a monitoring well, flowing bottom-hole pressure data at the injection well, and time-lapse pressure measurements at several locations along the monitoring well. The history matching results indicate that the CO2 movement is mostly restricted to the intended zones of injection which is consistent with an independent warm-back analysis of the temperature data. In addition to employing simulation models and inverse methods for CO2 plume imaging, we also initialized a data-driven technology for detecting inter-well connectivity based on production and pressure data. Our machine-learning framework is built on the statistical recurrent unit (SRU) model and interprets well-based injection/production data into inter-well connectivity without relying on a geologic model. We test it on synthetic and field-scale CO2 EOR projects utilizing the water-alternating-gas (WAG) process. The validation of the proposed data-driven inter-well connectivity assessment is performed using synthetic data from simulation models where inter-well connectivity can be easily measured using the streamline-based flux allocation. The SRU model is shown to offer excellent prediction performance on the synthetic case. Despite significant measurement noise and frequent well shut-ins imposed in the field-scale case, the SRU model offers good prediction accuracy, the overall relative error of the phase production rates at most producers ranges from 10% to 30%. It is shown that the dominant connections identified by the data-driven method and streamline method are in close agreement. Texas A&M University, the lead organization in the project, was primarily responsible for the development of tomographic approaches for CO2 plume mapping in conjunction with distributed pressure, temperature and seismic onset time data. Battelle, as a subcontractor, was primarily responsible for the development of analytical and empirical methods for analyzing transient injection rate and pressure data from point/line sources such as injection and monitoring wells. An additional area of emphasis for Battelle was the use of machine learning for such tasks as inferring reservoir connectivity information from injection-production data, and identifying variable importance for machine learning-based proxy models developed from full-physics simulations. The two organizations also collaborated on the application of the tomographic inversion methodology for a field data set.

02 PETROLEUM↗

A comparison of eight optimization methods applied to a wind farm layout optimization problem

Abstract. Selecting a wind farm layout optimization method is difficult. Comparisons between optimization methods in different papers can be uncertain due to the difficulty of exactly reproducing the objective function. Comparisons by just a few authors in one paper can be uncertain if the authors do not have experience using each algorithm. In this work we provide an algorithm comparison for a wind farm layout optimization case study between eight optimization methods applied, or directed, by researchers who developed those algorithms or who had other experience using them. We provided the objective function to each researcher to avoid ambiguity about relative performance due to a difference in objective function. While these comparisons are not perfect, we try to treat each algorithm more fairly by having researchers with experience using each algorithm apply each algorithm and by having a common objective function provided for analysis. The case study is from the International Energy Association (IEA) Wind Task 37, based on the Borssele III and IV wind farms with 81 turbines. Of particular interest in this case study is the presence of disconnected boundary regions and concave boundary features. The optimization methods studied represent a wide range of approaches, including gradient-free, gradient-based, and hybrid methods; discrete and continuous problem formulations; single-run and multi-start approaches; and mathematical and heuristic algorithms. We provide descriptions and references (where applicable) for each optimization method, as well as lists of pros and cons, to help readers determine an appropriate method for their use case. All the optimization methods perform similarly, with optimized wake loss values between 15.48 % and 15.70 % as compared to 17.28 % for the unoptimized provided layout. Each of the layouts found were different, but all layouts exhibited similar characteristics. Strong similarities across all the layouts include tightly packing wind turbines along the outer borders, loosely spacing turbines in the internal regions, and allocating similar numbers of turbines to each discrete boundary region. The best layout by annual energy production (AEP) was found using a new sequential allocation method, discrete exploration-based optimization (DEBO). Based on the results in this study, it appears that using an optimization algorithm can significantly improve wind farm performance, but there are many optimization methods that can perform well on the wind farm layout optimization problem, given that they are applied correctly.

17 WIND ENERGY↗

Efficient Parallelization of Irregular Applications on GPU Architectures

With the enlarging computation capacity of general Graphics Processing Units (GPUs), leveraging GPUs to accelerate parallel applications has become a critical topic in academia and industry. However, a wide range of irregular applications with the computation-/memory-intensive nature cannot easily achieve high GPU utilization. The challenges mainly involve the following aspects: first, data dependence leads to coarse-grained kernel and inefficient parallelism; second, heavy GPU memory usage may cause frequent memory evictions and extra overhead of I/O; third, specific computation patterns produce memory redundancies; last, workload balance and data reusability conjunctly benefit the overall performance, but there may exist a dynamic trade-off between them. Targeting these challenges, this dissertation proposes multiple optimizations to accelerate two real-world applications: many-body correlation functions to simulate nuclear physics in a large-scale scientific system; the other is the eALS-based matrix factorization recommendation system. To accelerate the calculations of many-body correlation functions, this dissertation presents three frameworks in GPU memory management and multi-GPU scheduling. Firstly, an optimized systematic GPU memory management framework, MemHC, utilizes a series of new memory reduction designs in GPU memory allocation, CPU/GPU communications, and GPU memory oversubscription. Secondly, an enhanced multi-GPU scheduling framework, MICCO, particularly by taking both data dimension (e.g., data reuse and data eviction) and computation dimension into account. MICCO designs a heuristic scheduling algorithm and a machine learning-based regression model to generate the optimal settings of a proposed new concept to manage the trade-off. Thirdly, a locality-aware multi-GPU scheduling framework. This scheduler leverages pipeline batch generation with a looking-ahead strategy by building local dependency graphs for memory transfer reduction and better data reuse, achieving up to 79.92% memory cost reduction and 1.67x speedup. To parallelize the eALS-based recommendation system, this dissertation proposes an efficient CPU/GPU heterogeneous recommendation system, HEALS. HEALS employs newly designed architecture-adaptive data formats to achieve load balance and good data locality on CPU and GPU. To mitigate the data dependence, HEALS presents a CPU/GPU collaboration model for both task parallelism and data parallelism with multiple kernel computation optimizations. In summary, this dissertation efficiently accelerates two typical irregular applications on GPUs by building four frameworks, including CPU/GPU collaboration, GPU memory management, and multi-GPU scheduling.

Wang, Qihan↗

Task 12 Sustainability - Methodological Guidelines on Net Energy Analysis of Photovoltaic Electricity (2nd Edition)

Net Energy Analysis (NEA) is a structured, comprehensive method of quantifying the extent to which a given energy source is able to provide a net energy gain (i.e., an energy surplus) to the end user, after accounting for all the energy losses occurring along the chain of processes that are required to exploit it (i.e., for its extraction, processing and transformation into a usable energy carrier, and delivery to the end user), as well as for all the additional energy 'investments' that are required in order to carry out the same chain of processes. However, this general framework leaves the individual practitioner with a range of choices that can affect the results and thus, the conclusions of a NEA study. The current IEA PVPS guidelines were developed to provide guidance on assuring consistency, balance, and quality to enhance the credibility and reliability of the results from photovoltaic (PV) NEAs. The guidelines represent a consensus among the authors - PV NEA experts in North America and Europe - for assumptions made on PV performance, process inputs and outputs, methods of analysis, and reporting of the results. Guidance is given on photovoltaic-specific parameters used as inputs in NEA and on choices and assumptions in inventory data analysis and on implementation of modelling approaches. A consistent approach towards system modelling, the functional unit, the system boundaries and allocation aspects enhance the credibility of PV electricity NEA studies and enables balanced NEA-based comparisons. Specifically, "apples-to-oranges" comparisons of different energy carriers (e.g., fuels vs. electricity) are not methodologically sound and are to be avoided in all cases; also, any comparison across renewable and non-renewable electricity generation technologies must clearly point out the intrinsically short-term nature of the NEA viewpoint, which does not capture the long-term sustainability implications of renewable vs. non-renewable primary energy harvesting and use: non-renewable primary energy resources are depleted and finally exhausted (irrespective of the size of the EROI), while renewable primary energy resources are not. This document provides an in-depth discussion of a common metric of NEA, namely the energy return on investment (EROI), and how this is to be interpreted vis-a-vis the deceptively similar-sounding metrics in the field of Life Cycle Assessment (LCA): cumulative energy demand (CED) and non-renewable cumulative energy demand (nr-CED) per unit output. Specifically, a number of key differences are highlighted between these metrics as applied to electricity production systems, which are listed in Table S-1.

14 SOLAR ENERGY↗

Tracking Volumetric Units in Modular Factories for Automated Progress Monitoring Using Computer Vision

The construction industry is increasingly adopting off-site and prefabricated methods due to advantages offered in safety, quality, and lead time. Applying industrialized methods for plant management in offsite construction factories requires the collection of large volumes of production process data, which is a tedious task when performed manually. Recent attempts to automate this process have relied on sensor-based data collection methods which are susceptible to noise, expensive, and difficult to validate. Computer vision methods, however, enable process data collection from videos without the limitations of the other sensor-based methods. This technology has not been applied for offsite construction except in very few instances and therefore, this study proposes a novel method to reliably collect the production process data using computer vision method in near real-time from widely used surveillance cameras in offsite construction. The proposed method allows the user to annotate the workstations of interest on the video as ground truths and process these areas throughout the entire video to track the units entering and leaving stations, while continuously updating a near real-time schedule of the production line. This framework was validated by implementing on the surveillance videos of the production process of modular home manufacturing in a factory. The results consistently provided 100% accuracy, after denoising, for all the videos processed including 60 h of work for a station. The developed method enables real-time tracking of station performance, which can enable continuous improvement methods for factory management and resource allocation.

computer vision↗

The HTAP_v3 emission mosaic: merging regional and global monthly emissions (2000–2018) to support air quality modelling and policies

This study, performed under the umbrella of the Task Force on Hemispheric Transport of Air Pollution (TF-HTAP), responds to the global and regional atmospheric modelling community's need of a mosaic emission inventory of air pollutants that conforms to specific requirements: global coverage, long time series, spatially distributed emissions with high time resolution, and a high sectoral resolution. The mosaic approach of integrating official regional emission inventories based on locally reported data, with a global inventory based on a globally consistent methodology, allows modellers to perform simulations of high scientific quality while also ensuring that the results remain relevant to policymakers. HTAP_v3, an ad hoc global mosaic of anthropogenic inventories, has been developed by integrating official inventories over specific areas (North America, Europe, Asia including Japan and South Korea) with the independent Emissions Database for Global Atmospheric Research (EDGAR) inventory for the remaining world regions. The results are spatially and temporally distributed emissions of SO 2 , NO $x$ , CO, non-methane volatile organic compounds (NMVOCs), NH 3 , PM 10 , PM 2.5 , black carbon (BC), and organic carbon (OC), with a spatial resolution of 0.1º × 0.1º and time intervals of months and years, covering the period 2000–2018. The emissions are further disaggregated into 16 anthropogenic emitting sectors. This paper describes the methodology applied to develop such an emission mosaic, reports on source allocation, differences among existing inventories, and best practices for the mosaic compilation. One of the key strengths of the HTAP_v3 emission mosaic is its temporal coverage, enabling the analysis of emission trends over the past 2 decades. The development of a global emission mosaic over such long time series represents a unique product for global air quality modelling and for better-informed policymaking, reflecting the community effort expended by the TF-HTAP to disentangle the complexity of transboundary transport of air pollution.

54 ENVIRONMENTAL SCIENCES↗

Developing an Automated Uncertainty Quantification Tool to Improve Watershed-Scale Predictions of Water and Nutrient Cycling

Managing the flow of water, nutrients, and contaminants in watersheds is vital to addressing pressing issues related to water scarcity, access to clean drinking water, energy production, resilience to natural and anthropogenic perturbations, and ecological restoration. Decisions about the management of watersheds critically depend on the accuracy with which the flow of water and chemicals through the watershed can be predicted by computer models. Prediction uncertainty can be reduced by matching the model to data, which are collected in the field at great expense. The contribution of watershed characterization data to reducing uncertainty of relevant model predictions can be evaluated in a so-called data-worth analysis, which provides transparent, quantitative metrics about a data set’s value for the support of relevant watershed management objectives. To achieve this goal, we developed a software package that implements the data-worth analysis approach for use with state-of-the-art watershed models. The purpose of the proposed data-worth analysis is to help decision-makers allocate resources for watershed characterization such that the uncertainty in model predictions can be significantly reduced, which leads to better, more effective management decisions. At the same time, watershed characterization costs can be reduced. The specific technical objectives of this SBIR/STTR Phase II project were to develop a framework and associated software toolsets that implement the uncertainty quantification and data-worth analysis approach for use with state-of-the-art watershed models. This goal was achieved by (A) developing a user-friendly, robust software package that is accessible to a wide audience, including watershed managers, policy-makers, and public stakeholders; (B) by demonstrating application of the prototype on several use cases that are representative of complex watershed management challenges spanning a range of scales and that consider different open-source, DOE-based codes and other modeling platforms; and (C) by gathering information about the needs and requirements from potential users to help guide future developments, ensuring that the final product will be commercially viable. The developed software consists of a graphical user interface that guides the user through a sequence of analysis steps, supported by toolsets that leverage state-of-the-art computational simulation-optimization capabilities. A prototype of the software runs on multiple platforms (PC, Mac, multi-processor Linux environment), is linked to diverse watershed simulators (e.g., ECOSYS, TOUGH2, TOUGHREACT, Amanzi-ATS), performs multiple analysis tasks (predictive simulations, sensitivity analysis, uncertainty analysis, automatic parameter estimation, and data-worth analysis, multicomponent geothermometry), and is readily extensible to include external simulators and analysis tools. The software is being commercialized and will be continually updated to address user needs.

58 GEOSCIENCES↗