Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “high throughput computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

FPGA Implementation of Stereo Disparity with High Throughput for Mobility Applications

High speed stereo vision can allow unmanned robotic systems to navigate safely in unstructured terrain, but the computational cost can exceed the capacity of typical embedded CPUs. In this paper, we describe an end-to-end stereo computation co-processing system optimized for fast throughput that has been implemented on a single Virtex 4 LX160 FPGA. This system is capable of operating on images from a 1024 x 768 3CCD (true RGB) camera pair at 15 Hz. Data enters the FPGA directly from the cameras via Camera Link and is rectified, pre-filtered and converted into a disparity image all within the FPGA, incurring no CPU load. Once complete, a rectified image and the final disparity image are read out over the PCI bus, for a bandwidth cost of 68 MB/sec. Within the FPGA there are 4 distinct algorithms: Camera Link capture, Bilinear rectification, Bilateral subtraction pre-filtering and the Sum of Absolute Difference (SAD) disparity. Each module will be described in brief along with the data flow and control logic for the system. The system has been successfully fielded upon the Carnegie Mellon University's National Robotics Engineering Center (NREC) Crusher system during extensive field trials in 2007 and 2008 and is being implemented for other surface mobility systems at JPL.

Random access memory↗

Active learning of polarizable nanoparticle phase diagrams for the guided design of triggerable self-assembling superlattices

Polarizable nanoparticles are of interest in materials science because of their rich and complex phase behavior that can be used to engineer nanostructured materials with long-range crystalline order. To understand and rationally navigate the design space of polarizable nanoparticles for self-assembling highly ordered superlattices, we developed a coarse-grained computational model to describe the nanoparticle-nanoparticle interactions in implicit solvent and employ the computationally efficient image method to model many-body polarization interactions. We conducted high-throughput virtual screening over a five-dimensional particle design space spanned by temperature, particle size, particle charge, particle dielectric, and solvent dielectric using enhanced sampling molecular dynamics calculations within an active learning framework to efficiently map out the regions of thermodynamic stability of the self-assembled aggregates. We validate our predictions in comparisons against small angle x-ray scattering measurements of gold nanoparticles surface functionalized with metal chalcogenide ligands. Lastly, we use our validated phase maps to computationally design switchable nanostructured materials capable of triggered assembly and disassembly as a function of temperature and solvent dielectric with potential applications as sensors, smart windows, optoelectronic devices, and in medical diagnostics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

High-throughput terahertz imaging: progress and challenges

Abstract Many exciting terahertz imaging applications, such as non-destructive evaluation, biomedical diagnosis, and security screening, have been historically limited in practical usage due to the raster-scanning requirement of imaging systems, which impose very low imaging speeds. However, recent advancements in terahertz imaging systems have greatly increased the imaging throughput and brought the promising potential of terahertz radiation from research laboratories closer to real-world applications. Here, we review the development of terahertz imaging technologies from both hardware and computational imaging perspectives. We introduce and compare different types of hardware enabling frequency-domain and time-domain imaging using various thermal, photon, and field image sensor arrays. We discuss how different imaging hardware and computational imaging algorithms provide opportunities for capturing time-of-flight, spectroscopic, phase, and intensity image data at high throughputs. Furthermore, the new prospects and challenges for the development of future high-throughput terahertz imaging systems are briefly introduced.

36 MATERIALS SCIENCE↗

VA EDH Advanced Software Pipeline Framework Report: Enhancing Automation and Scalability

The VA Environmental Determinants of Health (EDH) Advanced Software Pipeline Framework is designed to enhance the efficiency, scalability, and security of geospatial data processing workflows. This framework integrates modern data orchestration and containerization technologies, including Prefect for workflow automation, Docker for containerization, and PostgreSQL/PostGIS for geospatial data storage and analysis. It ensures standardized, reproducible, and automated data processing, supporting VA objectives related to substance use risk assessment and recovery research. The pipeline addresses key scalability and performance challenges through horizontal and vertical scaling, high-performance computing (HPC) integration, parallel processing, task caching, and dynamic resource allocation. These optimizations improve throughput and reduce latency, allowing the system to efficiently manage large and complex datasets. Additionally, security and compliance measures—such as data encryption (SSL), Role-Based Access Control (RBAC), and adherence to GDPR and HIPAA standards—safeguard sensitive information throughout data transmission and storage. A key implementation of this framework includes the automation of shelter list geolocation workflows, ensuring that up-to-date data is readily available for VA decision-making. Lessons learned from this project include the transition from in-memory processing to incremental storage writes, improving resource management and reliability. Future enhancements aim to expand automation, integrate AI-driven anomaly detection, and incorporate high-performance computing resources. This framework provides a scalable, secure, and adaptable solution for managing geospatial datasets, reinforcing the VA’s ability to support clinical and strategic initiatives through data-driven decision-making.

97 MATHEMATICS AND COMPUTING↗

Synchrotron‐source micro‐x‐ray computed tomography for examining butterfly eyes

Comparative anatomy is an important tool for investigating evolutionary relationships among species, but the lack of scalable imaging tools and stains for rapidly mapping the microscale anatomies of related species poses a major impediment to using comparative anatomy approaches for identifying evolutionary adaptations. We describe a method using synchrotron source micro-x-ray computed tomography (syn-μXCT) combined with machine learning algorithms for high-throughput imaging of Lepidoptera (i.e., butterfly and moth) eyes. Our pipeline allows for imaging at rates of ~15 min/mm 3 at 600 nm 3 resolution. Image contrast is generated using standard electron microscopy labeling approaches (e.g., osmium tetroxide) that unbiasedly labels all cellular membranes in a species-independent manner thus removing any barrier to imaging any species of interest. To demonstrate the power of the method, we analyzed the 3D morphologies of butterfly crystalline cones, a part of the visual system associated with acuity and sensitivity and found significant variation within six butterfly individuals. Despite this variation, a classic measure of optimization, the ratio of interommatidial angle to resolving power of ommatidia, largely agrees with early work on eye geometry across species. We show that this method can successfully be used to determine compound eye organization and crystalline cone morphology. Our novel pipeline provides for fast, scalable visualization and analysis of eye anatomies that can be applied to any arthropod species, enabling new questions about evolutionary adaptations of compound eyes and beyond.

59 BASIC BIOLOGICAL SCIENCES↗

Predicting Sintering Window of Binder Jet Additively Manufactured Parts Using a Coupled Data Analytics and CALPHAD Approach

Batch-to-batch variation in powder compositions for binder jet additive manufacturing (BJAM) can significantly deter defining an “ideal” sintering window for a given alloy. One way to overcome the problem is by running sintering experiments at various temperatures for each batch of the powder. However, such an approach increases the time required to achieve large-scale production of parts. The predictive capabilities of computational thermodynamic tools like CALPHAD can be leveraged to overcome the challenge, especially for binder jet additive manufacturing, since the process occurs under near-equilibrium conditions. However, calculating the sintering window using CALPHAD can be computationally expensive, considering many possible feedstock compositions within “specification”. Here, we generate high throughput CALPHAD data for nickel-based superalloys to develop machine learning models to predict the sintering window rapidly. The predictive capability of the models has been validated using published results on BJAM of Inconel 718 and 625. Further, validated models are lightweight and can be deployed in an industrial setting to get sintering window in an accelerated manner.

36 MATERIALS SCIENCE↗

Computational prediction of ferromagnetic 𝐴⁢𝑇 6 ⁢𝑋 6 kagome compounds

We present a systematic high-throughput density-functional-theory investigation of the structural and magnetic stability of 312 substitutional compounds in the magnetic kagome 𝐴⁢𝑇 6⁢ 𝑋 6 family. Our screening confirms the stability of many previously reported structures and predicts several additional stable candidates. Within collinear spin configurations, we find that Fe-based systems predominantly adopt antiferromagnetic ground states, whereas Mn-based analogs exhibit a more balanced distribution between ferromagnetic and antiferromagnetic order. For compounds exhibiting several nearly degenerate collinear configurations, we analyze the nature of their magnetic ground states, assess the possible emergence of noncollinear order, and discuss the limitations and uncertainties inherent to standard density-functional approaches. Our electronic structure analysis further reveals that predicted ferromagnetic kagome systems display characteristic features of topological metals, with rich magnetic configurations that can be tuned by chemical substitution. Altogether, these ferromagnetic kagome compounds constitute a broad and still largely unexplored materials platform for the emergence of exciting magnetotransport phenomena.

Density functional calculations↗

Characterizing the Impact of GPU Power Management on an Exascale System

As GPU-accelerated high-performance computing (HPC) systems approach exascale performance, controlling energy consumption without compromising throughput is essential. Architectures such as the AMD MI250X-based Frontier supercomputer provide runtime mechanisms like frequency and power capping, enabling energy tuning without modifying application code. Although both target energy reduction, they operate via distinct hardware control paths and influence workloads differently. We present a comprehensive evaluation of these strategies on a leadership-class system using diverse HPC proxy applications representative of production workloads. Our study analyzes performance–energy trade-offs across multiple capping levels, node counts (1 and 32), and application profiles. Results show that frequency capping generally achieves higher energy efficiency and scalability, with gains of up to 13.2% without performance loss, while power capping is more effective for single-node runs or bursty GPU utilization. We also provide practical guidelines to help system administrators and users balance energy efficiency and performance in large-scale scientific workloads.

Costa, Mariana [Universidade Federal do Rio Grande↗

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES↗

Heterogeneous Reconstruction of Tracks and Primary Vertices With the CMS Pixel Tracker

The High-Luminosity upgrade of the Large Hadron Collider (LHC) will see the accelerator reach an instantaneous luminosity of 7 × 10 34 cm −2 s −1 with an average pileup of 200 proton-proton collisions. These conditions will pose an unprecedented challenge to the online and offline reconstruction software developed by the experiments. The computational complexity will exceed by far the expected increase in processing power for conventional CPUs, demanding an alternative approach. Industry and High-Performance Computing (HPC) centers are successfully using heterogeneous computing platforms to achieve higher throughput and better energy efficiency by matching each job to the most appropriate architecture. In this paper we will describe the results of a heterogeneous implementation of pixel tracks and vertices reconstruction chain on Graphics Processing Units (GPUs). The framework has been designed and developed to be integrated in the CMS reconstruction software, CMSSW. The speed up achieved by leveraging GPUs allows for more complex algorithms to be executed, obtaining better physics output and a higher throughput.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

The Matsu Wheel: A Cloud-Based Framework for Efficient Analysis and Reanalysis of Earth Satellite Imagery

Project Matsu is a collaboration between the Open Commons Consortium and NASA focused on developing open source technology for cloud-based processing of Earth satellite imagery with practical applications to aid in natural disaster detection and relief. Project Matsu has developed an open source cloud-based infrastructure to process, analyze, and reanalyze large collections of hyperspectral satellite image data using OpenStack, Hadoop, MapReduce and related technologies. We describe a framework for efficient analysis of large amounts of data called the Matsu "Wheel." The Matsu Wheel is currently used to process incoming hyperspectral satellite data produced daily by NASA's Earth Observing-1 (EO-1) satellite. The framework allows batches of analytics, scanning for new data, to be applied to data as it flows in. In the Matsu Wheel, the data only need to be accessed and preprocessed once, regardless of the number or types of analytics, which can easily be slotted into the existing framework. The Matsu Wheel system provides a significantly more efficient use of computational resources over alternative methods when the data are large, have high-volume throughput, may require heavy preprocessing, and are typically used for many types of analysis. We also describe our preliminary Wheel analytics, including an anomaly detector for rare spectral signatures or thermal anomalies in hyperspectral data and a land cover classifier that can be used for water and flood detection. Each of these analytics can generate visual reports accessible via the web for the public and interested decision makers. The result products of the analytics are also made accessible through an Open Geospatial Compliant (OGC)-compliant Web Map Service (WMS) for further distribution. The Matsu Wheel allows many shared data services to be performed together to efficiently use resources for processing hyperspectral satellite image data and other, e.g., large environmental datasets that may be analyzed for many purposes.

Synthetic-domain computing and neural networks using lithium niobate integrated nonlinear phononics

Analogue computing uses the physical behaviours of devices to provide energy-efficient arithmetic operations. However, scaling up analogue computing platforms by simply increasing the number of devices leads to challenges such as device-to-device variation. Here, in this study, we report scalable analogue computing and neural networks in the synthetic frequency domain using an integrated nonlinear phononic platform on lithium niobate. This synthetic-domain computing is robust to device variations, as vectors and matrices are concurrently encoded at different frequencies within a single device, achieving a high throughput per area. Leveraging inherent nonlinearities, our device-aware neural network can perform a four-class classification task with an accuracy of 98.2%. The nonlinear phononic computing hardware also maintains consistent performance over a wide operational temperature range (characterized up to 192 °C). Our synthetic-domain computing combines single-device parallelism, inherent nonlinearity and environmental stability, and could be of use in edge computing applications in which power efficiency and environmental resilience are crucial.

Ji, Jun [Virginia Polytechnic Inst. and State Univ↗

Defect Diffusion Graph Neural Networks for Materials Discovery in High-Temperature Energy Applications

Here, the migration of crystallographic defects dictates material properties and performance for a plethora of technological applications. Density functional theory (DFT)-based nudged elastic band (NEB) calculations are a powerful computational technique for predicting defect migration activation energy barriers, yet they become prohibitively expensive for high-throughput screening of defect diffusivities. Without introducing hand-crafted (i.e., chemistry- or structure-specific) descriptors, we propose a generalized deep learning approach to train surrogate models for NEB energies of vacancy migration by hybridizing graph neural networks with transformer encoders and simply using pristine host structures as input. With sufficient training data, computationally efficient and simultaneous inference of vacancy defect thermodynamics and migration activation energies can be obtained to compute temperature-dependent vacancy diffusivities and to down-select candidates for more thorough DFT analysis or experiments. Thus, as we specifically demonstrate for potential water-splitting materials, candidates with desired defect thermodynamics, kinetics, and host stability properties can be more rapidly targeted from open-source databases of experimentally validated or hypothetical materials.

14 SOLAR ENERGY↗

A Bayesian nonparametric analysis for zero-inflated multivariate count data with application to microbiome study

High-throughput sequencing technology has enabled researchers to profile microbial communities from a variety of environments, but analysis of multivariate taxon count data remains challenging. Here, we develop a Bayesian nonparametric (BNP) regression model with zero inflation to analyse multivariate count data from microbiome studies. A BNP approach flexibly models microbial associations with covariates, such as environmental factors and clinical characteristics. The model produces estimates for probability distributions which relate microbial diversity and differential abundance to covariates, and facilitates community comparisons beyond those provided by simple statistical tests. We compare the model to simpler models and popular alternatives in simulation studies, showing, in addition to these additional community-level insights, it yields superior parameter estimates and model fit in various settings. The model's utility is demonstrated by applying it to a chronic wound microbiome data set and a Human Microbiome Project data set, where it is used to compare microbial communities present in different environments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

express: Extensible, high-level workflows for swifter ab initio materials modeling

In this work, we introduce an open-source Julia project, express, an extensible, lightweight, high-throughput, high-level workflow framework that aims to automate ab initio calculations for the materials science community. express is shipped with well-tested workflow templates, including structure optimization, equation of state (EOS) fitting, phonon spectrum (lattice dynamics) calculation, and thermodynamic property calculation in the framework of the quasi-harmonic approximation (QHA). It is designed to be highly modularized so that its components can be reused across various occasions, and customized workflows can be built on top of that. Users can also track the status of workflows in real-time, and rerun failed jobs thanks to the data lineage feature express provides. Finally, two working examples, i.e., all workflows applied to lime and akimotoite, are also presented in the code and this paper.

36 MATERIALS SCIENCE↗

Comparison of DeePMD, MTP, GAP, ACE and MACE Machine‐Learned Potentials for Radiation‐Damage Simulations: A User Perspective

Accurate and efficient interatomic potentials are essential for molecular dynamics (MD) simulations of radiation damage, gas diffusion, and phase stability in complex ceramics such as LiAlO 2 , especially under extreme conditions relevant to tritium production. Here, we evaluate the performance of six machine-learned interatomic potentials (MLIPs), moment tensor potential (MTP), Gaussian approximation potential, deep potential (DeePMD), atomic cluster expansion (ACE), message-passing ACE (multilayer atomic cluster expansion (MACE) pretrained) and MACE (trained from-scratch), all trained on the same density functional theory dataset with inclusion of tritium. The MLIPs are benchmarked against traditional Buckingham and ReaxFF potentials in terms of energy accuracy, density predictions, thermal equilibration behavior, threshold displacement energy (E d ), tritium diffusivity, and computational cost. Among the models, MTP shows the best overall balance between efficiency and accuracy, with low force and energy errors and realistic E d values for Li and Al. The ACE and MACE (pretrained and trained from scratch) models exhibit high E d (>200 eV) and unphysical pair interactions. DeePMD underestimates Ed due to overly repulsive behavior even at equilibrium distances. All models over-estimate tritium diffusion but the pretrained MACE model behaves well during tritium-diffusion simulations up to 500 K, maintaining diffusivities in the physically consistent 10 −11 m 2 /s range. Finally, we quantify the computational cost of each potential in large-scale atomic/molecular massively parallel simulator, finding that only MTP is more efficient than traditional empirical potentials, while others are significantly more expensive. These findings explain the trade-offs between accuracy and computational cost in MLIP development and provide essential guidance for use in high-throughput radiation damage and gas diffusion simulations in nuclear ceramics.

74 ATOMIC AND MOLECULAR PHYSICS↗

ABINIT: Overview and focus on selected capabilities

ABINIT is probably the first electronic-structure package to have been released under an open-source license about 20 years ago. It implements density functional theory, density-functional perturbation theory (DFPT), many-body perturbation theory (GW approximation and Bethe–Salpeter equation), and more specific or advanced formalisms, such as dynamical mean-field theory (DMFT) and the “temperature-dependent effective potential” approach for anharmonic effects. Relying on planewaves for the representation of wavefunctions, density, and other space-dependent quantities, with pseudopotentials or projector-augmented waves (PAWs), it is well suited for the study of periodic materials, although nanostructures and molecules can be treated with the supercell technique. The present article starts with a brief description of the project, a summary of the theories upon which ABINIT relies, and a list of the associated capabilities. It then focuses on selected capabilities that might not be present in the majority of electronic structure packages either among planewave codes or, in general, treatment of strongly correlated materials using DMFT; materials under finite electric fields; properties at nuclei (electric field gradient, Mössbauer shifts, and orbital magnetization); positron annihilation; Raman intensities and electro-optic effect; and DFPT calculations of response to strain perturbation (elastic constants and piezoelectricity), spatial dispersion (flexoelectricity), electronic mobility, temperature dependence of the gap, and spin-magnetic-field perturbation. The ABINIT DFPT implementation is very general, including systems with van der Waals interaction or with noncollinear magnetism. Community projects are also described: generation of pseudopotential and PAW datasets, high-throughput calculations (databases of phonon band structure, second-harmonic generation, and GW computations of bandgaps), and the library libpaw. ABINIT has strong links with many other software projects that are briefly mentioned.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

FLASH: FPGA-Accelerated Smart Switches with GCN Case Study

Some communication switches, e.g., the Mellanox SHArP and those in the IBM BlueGene clusters, are augmented to process packets at the application level with fixed-function collectives. This approach, however, lacks flexibility, which limits their applicability in diverse and dynamic workloads. Recently, a new type of programmable packet processor, which uses high-level languages, e.g., P4, has emerged as possible candidates. P4-based switches, however, fall short in certain applications, including machine learning, where capabilities not currently supported by P4 are needed. These include more complex calculation, such as sparse computation and fused multiply-accumulate, data-intensive floating point operations, data reuse, and significant memory. The problem addressed here is that such a switch augmentation needs to support: a large amount of state, significant flexible compute capability, and ease of programming, all while maintaining full functionality, including ensuring high throughput, and demonstrating utility. In this work, we propose a programmable look-aside-type accelerator that can be embedded into, or attached to, existing communication switch pipelines and that is capable of processing packets at line-rate. The proposed in-switch accelerator is based on mixing an ISA (subset of RISC-V instructions) with dataflow graphs (found in CGRAs). To augment performance, vector instructions are also supported. To facilitate usability, we have developed a complete toolchain to compile user-provided C/C++ codes to appropriate back-end instructions for configuring the accelerator. While this approach is flexible enough to support various workloads, in this paper, we consider Graph Convolutional Networks (GCNs) as a case study. Experimental results show that this approach considerably improves the performance of distributed GCN applications.

Haghi, Pouya↗