Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “high throughput computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Telemetry handling on the Space Station data management system

This paper examines the impact of telemetry handling on the design of the onboard networks that are part of the Space Station Data Management System (DMS). An architectural approach to satisfying the DMS requirement for support of the high throughput needed for telemetry transport and for servicing distributed computer systems is discussed. Several of the functionality vs. performance tradeoffs that must be made in developing an optimized mechanism for handling telemetry data in the DMS are considered.

Whitelaw, Virginia A.↗

Balanced-Load Real-Time Multiprocessor System

Modularity and parallelism provide tolerance to faults and high throughput capacity. System, called MAX, is network of interconnected computers. MAX cluster consists of group of modules - each semiautonomous computer. Modules connected to each other and to other clusters by global bus and circuit-switched communication mesh.

Rasmussen, Robert D.↗

Response of Subsurface Nitrogen-Cycling Microbial Communities to Environmental Fluctuations (Final Technical Report)

Riparian floodplains are dynamic ecosystems linking terrestrial and riverine systems. These floodplains experience hydrological shifts such as changes in water table height, flooding, and drought and can be ‘hotspots’ of biogeochemical cycling due to shifting sediment moisture (and saturation) and subsurface exchanges of water, nutrients, and other compounds across different sediment layers. Subsurface microbial communities are the primary drivers of biogeochemical processes in floodplains, and thus their structure and function can directly influence both surface and groundwater quality. The microbial nitrogen (N) cycle is particularly important in floodplains as it affects nutrient availability and removal. Two functional guilds of chemoautotrophic (i.e. CO2-fixing) microorganisms are responsible for the first oxidative step of the N cycle, nitrification: ammonia-oxidizing archaea (AOA) and bacteria (AOB) catalyze the oxidation of ammonia to nitrite, while nitrite-oxidizing bacteria (NOB) oxidize nitrite to nitrate. Despite the critical role nitrification plays in N-cycling in both terrestrial and aquatic ecosystems, our understanding of the diversity, ecophysiology, and activity of nitrifying organisms in subsurface floodplain soils/sediments is extremely limited. To help address this critical knowledge gap, the overarching goal of this project was to determine how shifts in key environmental parameters and gradients impact microbial N-cycling communities/processes, with particular emphasis on nitrification, within hydrologically-variable floodplain sediments in the Wind River Basin near Riverton, Wyoming. The three specific objectives of this project were to: (1) to associate in situ environmental drivers of N cycling with distinct functional guilds; (2) determine the guild response to variation in key ecosystem drivers; and (3) develop a dynamic ecosystem model of the microbial N cycle with the Riverton subsurface using community genomic and biogeochemical data collected in the first two objectives. Over the course of this project, we employed both 16S rRNA gene amplicon sequencing and genome-resolved metagenomics to examine the phylogenetic diversity and metabolic potential of subsurface nitrifier communities within 68 samples collected across multiple sites, depths, and time points within the Riverton floodplain, allowing for both spatial and temporal investigations at different scales. This project benefitted tremendously from recent advances in high-throughput sequencing technologies coupled with dramatic improvements in the computational tools and algorithms available for analyzing such large, complex genomic datasets. By pairing these cutting-edge genomic approaches with depth-resolved sampling and detailed geochemical analyses of the Riverton floodplain, we have gained novel insights into the structure and function of subsurface nitrifier communities in relation to both hydrology and biogeochemistry. This project resulted in the most detailed and comprehensive characterization of N-cycling floodplain microbial communities to date and will hopefully inspire and pave the way for future studies using similar approaches in other floodplains. Indeed, such information is critical for understanding subsurface biogeochemical cycling and how elemental stores are altered from perturbations initiated by the water cycle within floodplains. Finally, because of the terrestrial-aquatic nature of the Riverton floodplain, results from this project are also of relevance to disciplines such as soil science, estuarine science, limnology & oceanography, biogeochemistry, geobiology, environmental engineering, as well as genomics and data science.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning-Guided Identification of PET Hydrolases from Natural Diversity

The enzymatic depolymerization of poly(ethylene terephthalate) (PET) is emerging as a leading chemical recycling technology for waste polyester. As part of this endeavor, new candidate enzymes identified from natural diversity can serve as useful starting points for enzyme evolution and engineering. In this study, we improved upon HMM searches by applying an iterative machine learning strategy to identify 400 putative PET-degrading enzymes (PET hydrolases) from naturally occurring homologs. Using high-throughput (HTP) experimental techniques, we successfully expressed and purified >200 enzyme candidates and assayed them for PET hydrolysis activity as a function of pH, temperature, and substrate crystallinity. From this library, we discovered 91 previously unknown PET hydrolases, 35 of which retain activity at pH 4.5 on crystalline material, which are conditions relevant to developing more efficient commercial processes. Notably, four enzymes showed equal to or higher activity than LCC-ICCG, a benchmark PET hydrolase, at this challenging condition in our screening assay, and 11 of which have pH optima <7. Using these data, we identified regions of PETases statistically correlated to activity at lower pH. We additionally investigated the effect of condition-specific activity data on trained machine learning predictors and found a precision (putative hit rate) improvement of up to 30% compared to a Hidden Markov Model alone. Our findings show that by pointing enzyme discovery toward conditions of interest with multiple rounds of experimental and machine learning, we can discover large sets of active enzymes and explore factors associated with activity at those conditions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Picturing the future of food

Abstract High‐throughput phenotyping (HTP) has emerged as one of the most exciting and rapidly evolving spaces within plant science. The successful application of phenotyping technologies will facilitate increases in agricultural productivity. High‐throughput phenotyping research is interdisciplinary and may involve biologists, engineers, mathematicians, physicists, and computer scientists. Here we describe the need for additional interest in HTP and offer a primer for those looking to engage with the HTP community. This is a high‐level overview of HTP technologies and analysis methodologies, which highlights recent progress in applying HTP to foundational research, identification of biotic and abiotic stress, breeding and crop improvement, and commercial and production processes. We also point to the opportunities and challenges associated with incorporating HTP across food production to sustainably meet the current and future global food supply requirements.

59 BASIC BIOLOGICAL SCIENCES↗

Analysis of an optical imaging system prototype for autonomously monitoring zooplankton in an aquaculture facility

Traditional approaches to biomonitoring in aquatic systems, such as sample collection, sorting, and identification, require significant time and effort, thereby limiting the spatiotemporal resolution of sample collection. Additionally, collection and preservation of samples for subsequent taxonomic identification and enumeration leads to mortality of organisms. Recent advances in technologies that utilize optical imaging and machine learning have provided new opportunities to expedite biomonitoring and lead to significant cost savings. These technologies can be advantageous to scientists or managers that conduct routine biomonitoring to inform operations, as in the case of aquaculture facilities. The Small Aquatic Organism optical imaging system (SAO) is a high-throughput optical imaging and classification prototype system that relies on computer vision and machine learning (Support Vector Machines, or SVMs) to autonomously identify and enumerate aquatic organisms. The SAO provides a more sustainable method of collecting large volumes of data and has the benefit of being used in situ. In this study, we tested the performance of the SAO in providing comparable results to manual zooplankton community monitoring in ten ponds at an aquaculture facility. We performed a side-by-side study comparing the sampling methods of plankton tow nets, where major zooplankton taxonomic classes were manually identified and enumerated, to sampling with the SAO. Vouchered samples were used to develop a training library for the SAO, where classes consisted of water boatman and zooplankton groups: cladocerans, copepod adults, copepod nauplii, and rotifers. SAO imagery was manually classified and compared with predicted results for validation. Accuracy for the SVM classifier of the SAO was 37.4 %. Convolutional Neural Networks (CNN) and Random Forest classifiers were also applied to SAO imagery and image features for comparison. The best CNN model and our Random Forest model had accuracies of 80.4 % and 46.6 % respectively. Challenges faced included the small size of copepod nauplii and rotifers and the limited resolution of the imaging camera, although there are tradeoffs between imaging resolution and the sample processing rate. Furthermore, our comparison shows that advancement in both optical imaging and ML are needed in order for the SAO prototype to yield comparable results to manual community monitoring in an aquaculture facility.

54 ENVIRONMENTAL SCIENCES↗

A scalable framework for efficient coupling of thermal and microstructural simulations in additive manufacturing

Predicting microstructure evolution in metal additive manufacturing (AM) is important for process optimization, but spatiotemporal scale disparities between thermal transport and microstructure evolution create significant challenges for efficient data transfer between simulation codes. To address this, we present Stork, a scalable framework for coupling thermal and microstructural simulations. Stork uses a sparse data representation to identify and store active solidification sub-volumes, enabling highly parallel quad-linear interpolation from coarse thermal grids to fine microstructure grids without large intermediate storage. We demonstrate the framework by coupling the semi-analytic heat transfer code 3DThesis with the time-parallel cellular automata code Toucan. This approach achieves over two orders of magnitude reduction in data generation time and file size compared to prior workflows. Numerical studies show that quad-linear interpolation preserves grain morphology and crystallographic texture in laser powder bed fusion (LPBF) simulations for coarsening ratios up to 16. Overall, Stork provides a scalable pathway for high-throughput, component-scale AM simulations on modern high-performance computing systems.

36 MATERIALS SCIENCE↗

Benchmarking Density Functional Theory Methods for Efficient Calculations of a Strongly Correlated Li 1– x Ni 1– y O 2−δ System

Transition metal oxides (TMOs), such as LiNiO 2 , are promising candidates for energy storage and electronic devices due to their unique electronic properties, exceptional physical and chemical characteristics, and ability to adopt multiple oxidation states. However, accurately predicting their properties using mean-field density functional theory (DFT) is challenging due to the presence of strongly correlated d-electrons and the complex interplay between their structural, electronic, and magnetic responses. These challenges are further exacerbated by the need to model defects, surfaces, and interfaces, which require computationally efficient, large-scale simulations. To address these issues, we carry out a benchmark study on the Li 1–x NiO 2 system, evaluating the performance of several popular functionals. Our findings demonstrate that combining SCAN functional relaxation with single-step HSE calculations provides a practical and scalable computational strategy. This approach balances accuracy and efficiency, enabling high-throughput simulations of strongly correlated TMOs and improved predictive modeling capability of TMOs for practical applications.

25 ENERGY STORAGE↗

GWAS identifies candidate genes controlling adventitious rooting in Populus trichocarpa

Adventitious rooting (AR) is critical to the propagation, breeding, and genetic engineering of trees. The capacity for plants to undergo this process is highly heritable and of a polygenic nature; however, the basis of its genetic variation is largely uncharacterized. To identify genetic regulators of AR, we performed a genome-wide association study (GWAS) using 1148 genotypes of Populus trichocarpa. GWASs are often limited by the abilities of researchers to collect precise phenotype data on a high-throughput scale; to help overcome this limitation, we developed a computer vision system to measure an array of traits related to adventitious root development in poplar, including temporal measures of lateral and basal root length and area. GWAS was performed using multiple methods and significance thresholds to handle non-normal phenotype statistics and to gain statistical power. These analyses yielded a total of 277 unique associations, suggesting that genes that control rooting include regulators of hormone signaling, cell division and structure, reactive oxygen species signaling, and other processes with known roles in root development. Numerous genes with uncharacterized functions and/or cryptic roles were also identified. These candidates provide targets for functional analysis, including physiological and epistatic analyses, to better characterize the complex polygenic regulation of AR.

59 BASIC BIOLOGICAL SCIENCES↗

MARBLE: A Multi-GPU Aware Job Scheduler for Deep Learning on HPC Systems

Deep learning (DL) has become a key tool for solving complex scientific problems. However, managing the multi-dimensional large-scale data associated with DL, especially atop extant multiple graphics processing units (GPUs) in modern supercomputers poses significant challenges. Moreover, the latest high-performance computing (HPC) architectures bring different performance trends in training throughput compared to the existing studies. Existing DL optimizations such as larger batch size and GPU locality-aware scheduling have little effect on improving DL training throughput performance due to fast CPU-to-GPU connections. Additionally, DL training on multiple GPUs scales sublinearly. Thus, simply adding more GPUs to a system is ineffective. To this end, we design MARBLE, a first-of-its-kind job scheduler, which considers the non-linear scalability of GPUs at the intra-node level to schedule an appropriate number of GPUs per node for a job. By sharing the GPU resources on a node with multiple DL jobs, MARBLE avoids low GPU utilization in current multi-GPU DL training on HPC systems. Our comprehensive evaluation in the Summit supercomputer shows that MARBLE is able to improve DL training performance by up to 48.3% compared to the popular Platform Load Sharing Facility (LSF) scheduler. Compared to the state-of-the-art of DL scheduler, Optimus, MARBLE reduces the job completion time by up to 47%.

Han, Jingoo↗

Studying CPU and memory utilization of applications on Fujitsu A64FX and Nvidia Grace Superchip

ARM-based manycore CPU architectures are well-positioned to provide the rising memory throughput requirements of modern data intensive scientific applications in High Performance Computing (HPC). The Fujitsu A64FX CPU platform is based on the ARM v8.2A architecture, and is the processor of the flagship Japanese supercomputer - "Fugaku", which was previously ranked as the #1 supercomputer in the world according to the Top500 list. The Nvidia Grace superchip features 144 Neoverse V2 cores based on the ARMv9 architecture with 4x128b SVE2, providing exceptional computational power. The chip supports up to 480GB of memory, making it ideal for AI, machine learning, and scientific computing workloads. In this paper, we conduct a thorough performance exploration of a variety of parallel bandwidth-sensitive benchmarks and applications compiled with the native Fujitsu compiler on a Fugaku A64FX compute node and ARM (LLVM) Compiler on an NVIDIA Grace superchip compute node, engaging all the computational cores per cluster using OpenMP multithreading (assuming the cores can drive the available bandwidth). Our ultimate goals are to study the resource utilization of scientific applications and benchmarks on A64FX and Grace superchip, considering graph application scenarios ( GAP Benchmark suite) and eleven appli- cation proxies from the Rodinia heterogeneous benchmark suite (considering domains such as Data Mining, Bioinformatics, Fluid Dynamics, Pattern Recognition, etc.). Through exhaustive performance monitoring, we quantify the resource utilization of diverse OpenMP-based HPC applications on both the Fujitsu A64FX and the Nvidia Grace Superchip platforms.

benchmarking, Performance Analysis, High performan↗

Adaptive Computing (AC) [SWR-24-106]

The Adaptive Computing (AC) software stack supports goal-based computing, for which a simulation workload is created on the fly adapting to the results of calculations. Application-specific code defines an objective, which may be to solve an optimization problem or to train a surrogate model with minimal uncertainty. Then, the AC driver decides where in the design parameter space to run simulations to best achieve that objective. This process is iterative and online; as new data is returned from simulations, the AC driver chooses new simulations to run. The AC driver can strategically run simulations on distributed hardware resources (including high performance computing machines, cloud resources, and edge devices) to maximize throughput and obey resource constraints.

Griffin, Kevin [National Renewable Energy Laborato↗

Adaptive Computing (AC) (Open Source) [SWR-24-106]

The Adaptive Computing (AC) software stack supports goal-based computing, for which a simulation workload is created on the fly, adapting to the results of calculations. Application-specific code defines an objective, which may be to solve an optimization problem or to train a surrogate model with minimal uncertainty. Then, the AC driver decides where in the design parameter space to run simulations to best achieve that objective. This process is iterative and online; as new data is returned from simulations, the AC driver chooses new simulations to run. The AC driver can strategically run simulations on distributed hardware resources (including high performance computing machines, cloud resources, and edge devices) to maximize throughput and obey resource constraints.

Griffin, Kevin [National Laboratory of the Rockies↗

3D reconstruction identifies loci linked to variation in angle of individual sorghum leaves

Selection for yield at high planting density has reshaped the leaf canopy of maize, improving photosynthetic productivity in high density settings. Further optimization of canopy architecture may be possible. However, measuring leaf angles, the widely studied component trait of leaf canopy architecture, by hand is a labor and time intensive process. Here, we use multiple, calibrated, 2D images to reconstruct the 3D geometry of individual sorghum plants using a voxel carving based algorithm. Automatic skeletonization and segmentation of these 3D geometries enable quantification of the angle of each leaf for each plant. The resulting measurements are both heritable and correlated with manually collected leaf angles. This automated and scaleable reconstruction approach was employed to measure leaf-by-leaf angles for a population of 366 sorghum plants at multiple time points, resulting in 971 successful reconstructions and 3,376 leaf angle measurements from individual leaves. A genome wide association study conducted using aggregated leaf angle data identified a known large effect leaf angle gene, several previously identified leaf angle QTL from a sorghum NAM population, and novel signals. Genome wide association studies conducted separately for three individual sorghum leaves identified a number of the same signals, a previously unreported signal shared across multiple leaves, and signals near the sorghum orthologs of two maize genes known to influence leaf angle. Automated measurement of individual leaves and mapping variants associated with leaf angle reduce the barriers to engineering ideal canopy architectures in sorghum and other grain crops.

3D reconstruction↗

Graphics Processing Unit Assisted Thermographic Compositing

Objective: To develop a software application utilizing general purpose graphics processing units (GPUs) for the analysis of large sets of thermographic data. Background: Over the past few years, an increasing effort among scientists and engineers to utilize the GPU in a more general purpose fashion is allowing for supercomputer level results at individual workstations. As data sets grow, the methods to work them grow at an equal, and often great, pace. Certain common computations can take advantage of the massively parallel and optimized hardware constructs of the GPU to allow for throughput that was previously reserved for compute clusters. These common computations have high degrees of data parallelism, that is, they are the same computation applied to a large set of data where the result does not depend on other data elements. Signal (image) processing is one area were GPUs are being used to greatly increase the performance of certain algorithms and analysis techniques. Technical Methodology/Approach: Apply massively parallel algorithms and data structures to the specific analysis requirements presented when working with thermographic data sets.

Ragasa, Scott↗

Graphics Processing Unit Assisted Thermographic Compositing

Objective: To develop a software application utilizing general purpose graphics processing units (GPUs) for the analysis of large sets of thermographic data. Background: Over the past few years, an increasing effort among scientists and engineers to utilize the GPU in a more general purpose fashion is allowing for supercomputer level results at individual workstations. As data sets grow, the methods to work them grow at an equal, and often greater, pace. Certain common computations can take advantage of the massively parallel and optimized hardware constructs of the GPU to allow for throughput that was previously reserved for compute clusters. These common computations have high degrees of data parallelism, that is, they are the same computation applied to a large set of data where the result does not depend on other data elements. Signal (image) processing is one area were GPUs are being used to greatly increase the performance of certain algorithms and analysis techniques.

Ragasa, Scott↗

The Profiled Feldman-Cousins Method for Confidence Interval Construction for the Nova 3-Flavor Oscillation Analysis

The small interaction cross-section of neutrinos makes experimental neutrino physics particularly responsive to technological advancements. A significant development leveraged by the NOvA experiment is large-scale parallel processing, enabling novel computational approaches to longstanding experimental challenges. Central to managing the resulting high-throughput data is NOvA’s implementation of the Freight Train model, designed for efficient data production and handling.This dissertation details the methodology and execution of the NOvA 2024 3-Flavor Oscillation Analysis, supported by a comprehensive dataset spanning ten years. It emphasizes frequentist results refined through the Feldman-Cousins (FC) technique, specifically addressing confidence interval corrections in parameter estimation. The computational intensity associated with Feldman-Cousins arises from extensive Monte Carlo simulations, which were substantially mitigated through parallel computing on the Perlmutter supercomputer at the National Energy Research Scientific Computing Center (NERSC), employing the MPI framework.To further enhance computational efficiency, an Importance Sampling method is introduced and evaluated, demonstrating significant potential to reduce complexity, particularly in exploring extreme parameter space regions. This thesis presents both the successful application of advanced computational resources and the development of sophisticated statistical techniques, aiming to enhance the precision and scope of neutrino oscillation analyses.

Dye ajdye11190@gmail.com, Andrew Joseph [Mississip↗

Physics constrained unsupervised deep learning for rapid, high resolution scanning coherent diffraction reconstruction

By circumventing the resolution limitations of optics, coherent diffractive imaging (CDI) and ptychography are making their way into scientific fields ranging from X-ray imaging to astronomy. Yet, the need for time consuming iterative phase recovery hampers real-time imaging. While supervised deep learning strategies have increased reconstruction speed, they sacrifice image quality. Furthermore, these methods’ demand for extensive labeled training data is experimentally burdensome. Here, we propose an unsupervised physics-informed neural network reconstruction method, PtychoPINN, that retains the factor of 100-to-1000 speedup of deep learning-based reconstruction while improving reconstruction quality by combining the diffraction forward map with real-space constraints from overlapping measurements. In particular, PtychoPINN gains a factor of 4 in linear resolution and an 8 dB improvement in PSNR while also accruing improvements in generalizability and robustness. This blend of performance and computational efficiency offers exciting prospects for high-resolution real-time imaging in high-throughput environments such as X-ray free electron lasers (XFELs) and diffraction-limited light sources.

97 MATHEMATICS AND COMPUTING↗