Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “high throughput computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Computational thermodynamic study of SiC chemical vapor deposition from MTS-H 2

This study focuses on the computational thermodynamic analysis of the chemical vapor deposition (CVD) of SiC from the methyltrichlorosilane-hydrogen (MTS-H 2 ) using up-to-date thermodynamic databases. High-resolution computation has been performed with the fine intervals of temperature and pressure at the various H 2 /MTS ratios of interest to systematically investigate the deposition condition range (800 to 1600°C, 0 to 26 664 Pa, and H 2 /MTS ratios of 0.1 to 100) to guide experimental exploration. The influence of deposition parameters on the compositions and phase stabilities of the deposit and gas phase pertinent to vapor processing is elucidated. Low pressure and medium temperatures (1000 to 1400°C) are beneficial to reaching a higher SiC deposition efficiency and provide an optimal window for preparing a high-purity (>99 wt.% SiC) deposit. This optimal processing window expands significantly with an increasing H 2 /MTS ratio (<20). These results are supported by a number of previous theoretical and experimental observations. The mass fraction of SiC in deposit is proposed as an additional perspective to understand the discrepancy between thermodynamic calculation and experimental observation of pure CVD SiC at low H 2 /MTS ratios.

36 MATERIALS SCIENCE↗

Application of Systems Engineering Principles and Techniques in Biological Big Data Analytics: A Review

In the past few decades, we have witnessed tremendous advancements in biology, life sciences and healthcare. These advancements are due in no small part to the big data made available by various high-throughput technologies, the ever-advancing computing power, and the algorithmic advancements in machine learning. Specifically, big data analytics such as statistical and machine learning has become an essential tool in these rapidly developing fields. As a result, the subject has drawn increased attention and many review papers have been published in just the past few years on the subject. Different from all existing reviews, this work focuses on the application of systems, engineering principles and techniques in addressing some of the common challenges in big data analytics for biological, biomedical and healthcare applications. Specifically, this review focuses on the following three key areas in biological big data analytics where systems engineering principles and techniques have been playing important roles: the principle of parsimony in addressing overfitting, the dynamic analysis of biological data, and the role of domain knowledge in biological data analytics.

dynamic analysis↗

High-throughput native mass spectrometry as experimental validation for in silico drug design

In this project, we developed automated workflows for both experimental validation and computational prediction of protein-ligand interactions. The ultimate goal is to establish an integrated pipeline for high throughput design of inhibitors to enzymes relevant to all areas of biological research. Our experimental approach is based on native mass spectrometry (native MS), which measures accurate masses and quantify the relative abundance of protein-ligand complexes to define binding affinity. We set up an in-house built autosampler with highly flexible configurations to minimize the manual steps for high throughput native MS. In parallel, we also performed manual native MS to characterize the binding of substrates and inhibitors of SARS-Cov-2 nonstructural protein nsp10/16 in order to optimize the experimental parameters for future automation. On the computational side, we streamlined the pipeline to achieve minimal manual intervention for predicting enzyme inhibitors via simulation, using the same nsp10/16 system as an example. Using the native MS method we examined 8 top-ranked designed compounds, 2 of which showed weak binding of ~50 µM. The information from native MS experiment provided critical insights and the foundation for a fully integrated workflow for enzyme inhibitor design.

59 BASIC BIOLOGICAL SCIENCES↗

Density Functional Tight-Binding Models for Band Structures of Transition-Metal Alloys and Surfaces across the d -Block

First-principles electronic structure simulations are an invaluable tool for understanding chemical bonding and reactions. While machine-learning models such as interatomic potentials significantly accelerate the exploration of potential energy surfaces, electronic structure information is generally lost. Particularly in the field of heterogeneous catalysis, simulated electron band structures provide fundamental insights into catalytic reactivity. This ab initio knowledge is preserved in semiempirical methods such as density functional tight binding (DFTB), which extend the accessible computational length and time scales beyond first-principles approaches. In this paper here we present Shell-Optimized Atomic Confinement (SOAC) DFTB electronic-part-only parametrizations for bulk and surface band structures of all d-block transition metals that enable efficient predictions of electronic descriptors for large structures or high-throughput studies on complex systems outside the computational reach of density functional theory.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Optimal dimensionality selection for independent component analysis of transcriptomic data

Independent component analysis is an unsupervised machine learning algorithm that separates a set of mixed signals into a set of statistically independent source signals. Applied to high-quality gene expression datasets, independent component analysis effectively reveals both the source signals of the transcriptome as co-regulated gene sets, and the activity levels of the underlying regulators across diverse experimental conditions. Two major variables that affect the final gene sets are the diversity of the expression profiles contained in the underlying data, and the user-defined number of independent components, or dimensionality, to compute. Availability of high-quality transcriptomic datasets has grown exponentially as high-throughput technologies have advanced; however, optimal dimensionality selection remains an open question. We computed independent components across a range of dimensionalities for four gene expression datasets with varying dimensions (both in terms of number of genes and number of samples). We computed the correlation between independent components across different dimensionalities to understand how the overall structure evolves as the number of user-defined components increases. We then measured how well the resulting gene clusters reflected known regulatory mechanisms, and developed a set of metrics to assess the accuracy of the decomposition at a given dimension. We found that over-decomposition results in many independent components dominated by a single gene, whereas under-decomposition results in independent components that poorly capture the known regulatory structure. From these results, we developed a new method, called OptICA, for finding the optimal dimensionality that controls for both over- and under-decomposition. Specifically, OptICA selects the highest dimension that produces a low number of components that are dominated by a single gene. We show that OptICA outperforms two previously proposed methods for selecting the number of independent components across four transcriptomic databases of varying sizes. OptICA avoids both over-decomposition and under-decomposition of transcriptomic datasets resulting in the best representation of the organism’s underlying transcriptional regulatory network.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

High-Throughput Electric-Field-Assisted Sintering and Characterization Techniques for Materials Discovery

Despite improvements in computing and modeling capabilities, the performance of new materials, particularly those which deviate greatly in composition from well-studied materials (e.g., high-entropy alloys), can be difficult to simulate given the lack of available experimental property data. While some modeling techniques may attempt to predict the properties of these exotic materials, most are forced to make extrapolations from more traditional materials. To fulfill the need for accelerated material synthesis and property measurement, a high-throughput methodology has been developed. Utilizing electric-field-assisted sintering (EFAS), also known as spark plasma sintering (SPS), equipped with custom tooling, samples of differing alloy compositions can be produced simultaneously as a single alloy array. Several arrays have been produced with compositions spanning the Co-Cr-Fe-Mn-Ni alloy family, including many high-entropy alloys, while the novel array geometry has enabled the samples to be polished and characterized in parallel, using X-ray diffraction, scanning-electron microscopy, and laser-based thermal diffusivity measurements.

36 MATERIALS SCIENCE↗

Optimal decision-making in high-throughput virtual screening pipelines

Screening large pools of molecular candidates to identify those with specific design criteria or targeted properties is demanding in various science and engineering domains. While a high-throughput virtual screening (HTVS) pipeline can provide efficient means to achieving this goal, its design and operation often rely on experts' intuition, potentially resulting in suboptimal performance. In this paper, we fill this critical gap by presenting a systematic framework that can maximize the return on computational investment (ROCI) of such HTVS campaigns. Based on various scenarios, we empirically validate the proposed framework and demonstrate its potential to accelerate scientific discoveries through optimal computational campaigns, especially in the context of virtual screening.

97 MATHEMATICS AND COMPUTING↗

The Role of Cation Coordination in the Electrical and Optical Properties of Amorphous Transparent Conducting Oxides

Amorphous oxide semiconductor materials have demonstrated numerous advantages without compromise of electrical properties as compared to their crystalline counterparts, yet understanding of the fundamental principles allowing this has remained elusive. To study the origins of enhanced optoelectronic properties, we apply high-throughput, combinatorial sputtering, structural and spectral mapping, and computationally intensive ab initio molecular dynamics simulations with density functional theory to a ternary, post-transition metal oxide system, namely, zinc tin oxide. The deposited thin films exhibit a high figure of merit, achieving carrier densities in the range of 1019 to 1020 cm–3 and carrier mobilities up to 35 cm2/Vs. These results highlight the role of local distortions and cation coordination in determining the microscopic origins of carrier generation and transport. In particular, we identify the strong likelihood of Sn undercoordination in both Zn-poor and Zn-rich phases leading to the high carrier concentrations observed. This not only diverges from the still widespread historical indictment of oxygen vacancies controlling carrier population in crystalline oxides but also provides a comprehensive framework to describe the unique structure–property relationships using specific structural and electronic descriptors in disordered phase materials.

CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSICS,M↗

ESnet/JLab FPGA Accelerated Transport

To increase the science rate for high data rates/volumes, Thomas Jefferson National Accelerator Facility (JLab) has partnered with Energy Sciences Network (ESnet) to define an edge to data center traffic shaping / steering transport capability featuring data event aware network shaping and forwarding. The keystone of this ESnet+JLab FPGA Accelerated Transport (EJFAT) is the joint development of an AI/ML directed dynamic compute work Load Balancer (LB) of UDP streamed data. The LB is a suite consisting of a Field Programmable Gate Array (FPGA) executing the dynamically configurable, low fixed latency LB data plane featuring real-time packet redirection and high throughput, and a control plane running on the FPGA host computer that monitors network and compute farm telemetry in order to make dynamic AI/ML guided decisions for destination compute host redirection/load balancing and destination resource provisioning. The LB provides for three-tier horizontal scaling across LB suites, core compute hosts, and CPUs within a host. The LB effectively provides seamless integration of edge/core computing to support direct experimental data processing for immediate use by JLab science programs and others such as the EIC as well as data centers of the future requiring high throughput and low latency for both hot and cooled data for both running experiment data acquisition systems and data center use cases.

97 MATHEMATICS AND COMPUTING↗

Scaled-up Neuromorphic Array Communications Controller (SNACC) for Large-scale Neural Networks

Neuromorphic computing is one promising post-Moore’s law era technology, which takes inspiration from biological brains to perform computing tasks. The human brain contains billions of neurons with trillions of synapses and as neuromorphic hardware systems scale to larger and larger sizes, the communication system used to transfer information between neuromorphic elements and traditional computers must scale to keep up. In prior work, we describe the use of a separate neuromorphic array communications controller to support low-latency, high-throughput communication between our neuromorphic systems and a traditional computer. In this work, the neuromorphic array communications controller is used to support the scaling of a neuromorphic development system which uses multiple neuromorphic processors arranged in a two-dimensional array. The neuromorphic array communications controller, along with scalable local connections, is used to create a scalable neuromorphic platform to enable the development and testing of large neuromorphic network arrays.

Young, Aaron↗

Machine-Learning X-Ray Absorption Spectra to Quantitative Accuracy

Simulations of excited state properties, such as spectral functions, are often computationally expensive and therefore not suitable for high-throughput modeling. As a proof of principle, here we demonstrate that graph-based neural networks can be used to predict the x-ray absorption near-edge structure spectra of molecules to quantitative accuracy. Specifically, the predicted spectra reproduce nearly all prominent peaks, with 90% of the predicted peak locations within 1 eV of the ground truth. Besides its own utility in spectral analysis and structure inference, our method can be combined with structure search algorithms to enable high-throughput spectrum sampling of the vast material configuration space, which opens up new pathways to material design and discovery.

97 MATHEMATICS AND COMPUTING↗

Next-generation yeast-two-hybrid analysis with Y2H-SCORES identifies novel interactors of the MLA immune receptor

Protein-protein interaction networks are one of the most effective representations of cellular behavior. In order to build these models, high-throughput techniques are required. Next-generation interaction screening (NGIS) protocols that combine yeast two-hybrid (Y2H) with deep sequencing are promising approaches to generate interactome networks in any organism. However, challenges remain to mining reliable information from these screens and thus, limit its broader implementation. Here, we present a computational framework, designated Y2H-SCORES, for analyzing high-throughput Y2H screens. Y2H-SCORES considers key aspects of NGIS experimental design and important characteristics of the resulting data that distinguish it from RNA-seq expression datasets. Three quantitative ranking scores were implemented to identify interacting partners, comprising: 1) significant enrichment under selection for positive interactions, 2) degree of interaction specificity among multi-bait comparisons, and 3) selection of in-frame interactors. Using simulation and an empirical dataset, we provide a quantitative assessment to predict interacting partners under a wide range of experimental scenarios, facilitating independent confirmation by one-to-one bait-prey tests. Simulation of Y2H-NGIS enabled us to identify conditions that maximize detection of true interactors, which can be achieved with protocols such as prey library normalization, maintenance of larger culture volumes and replication of experimental treatments. Y2H-SCORES can be implemented in different yeast-based interaction screenings, with an equivalent or superior performance than existing methods. Proof-of-concept was demonstrated by discovery and validation of novel interactions between the barley nucleotide-binding leucine-rich repeat (NLR) immune receptor MLA6, and fourteen proteins, including those that function in signaling, transcriptional regulation, and intracellular trafficking.

59 BASIC BIOLOGICAL SCIENCES↗

Computationally Guided and Experimentally Validated Design of Custom Chelators for Critical Mineral Recovery

Selective, high throughput separation of target critical metals from complex environments such as fly ash leachates and mining process streams presents a significant challenge for economical production. Custom chelators and sorbents are an attractive technology for selective metal extraction, however it can be difficult to predict their performance, and significant experimental efforts are often required to develop chelating technologies. Here, we present a computational strategy focused on modelling chelator-metal binding interactions and benchmark these results versus experimental data. A computational pipeline combining forcefield, semiempirical, and meta-GGA methods with a thermodynamic framework optimized for error cancellation has been developed to predict binding energies of chelator complexes towards critical mineral recovery applications. This approach, originally validated on [2.2.2] cryptates binding mono- and divalent cations, demonstrated robust predictive capabilities with an R2 of 0.850 against experimental aqueous binding energies. The workflow includes metadynamics for exploring high-dimensional potential energy surfaces and a cluster-continuum model for accurate yet computationally efficient solvation modeling. Error cancellation between solvation energies of free and chelator-coordinated ions enables faster convergence, even with finite cluster sizes. Initial studies on the cryptates revealed consistent metal-ligand coordination patterns, with systematic variations influenced by ion size and charge, highlighting key structural features linked to binding selectivity. Further studies of a proprietary chelator have resulted in identification of previously unreported selectivity towards economically significant metals, which in-house experiments have confirmed, demonstrating the feasibility of this approach. By applying this methodology to new chelators targeting critical minerals such as lithium, cobalt, nickel and other strategic metals, we aim to accelerate the discovery of next-generation chelators for efficient recovery, recycling, and separation processes. This computational framework serves as the backbone of a high-throughput design pipeline tailored for sustainable resource utilization and may be applied to a wide range of systems to meet experimental needs.

computational materials↗

Reconstruction of Charged Particle Tracks in Realistic Detector Geometry Using a Vectorized and Parallelized Kalman Filter Algorithm

One of the most computationally challenging problems expected for the High-Luminosity Large Hadron Collider (HL-LHC) is finding and fitting particle tracks during event reconstruction. Algorithms used at the LHC today rely on Kalman filtering, which builds physical trajectories incrementally while incorporating material e ects and error estimation. Recognizing the need for faster computational throughput, we have adapted Kalman-filterbased methods for highly parallel, many-core SIMD and SIMT architectures that are now prevalent in high-performance hardware. Previously we observed significant parallel speedups, with physics performance comparable to CMS standard tracking, on Intel Xeon, Intel Xeon Phi, and (to a limited extent) NVIDIA GPUs. While early tests were based on artificial events occurring inside an idealized barrel detector, we showed subsequently that our mkFit software builds tracks successfully from complex simulated events (including detector pileup) occurring inside a geometrically accurate representation of the CMS-2017 tracker. Here, we report on advances in both the computational and physics performance of mkFit, as well as progress toward integration with CMS production software. Recently we have improved the overall eciency of the algorithm by preserving short track candidates at a relatively early stage rather than attempting to extend them over many layers. Moreover, mkFit formerly produced an excess of duplicate tracks; these are now explicitly removed in an additional processing step. We demonstrate that with these enhancements, mkFit becomes a suitable choice for the first iteration of CMS tracking, and eventually for later iterations as well. We plan to test this capability in the CMS High Level Trigger during Run 3 of the LHC, with an ultimate goal of using it in both the CMS HLT and oine reconstruction for the HL-LHC CMS tracker.

Cerati, Giuseppe↗

GPU-Accelerated Drug Discovery with Docking on the Summit Supercomputer: Porting, Optimization, and Application to COVID-19 Research

Protein-ligand docking is an in silico tool used to screen potential drug compounds for their ability to bind to a given protein receptor within a drug-discovery campaign. Experimental drug screening is expensive and time consuming, and it is desirable to carry out large scale docking calculations in a high-throughput manner to narrow the experimental search space. Few of the existing computational docking tools were designed with high performance computing in mind. Therefore, optimizations to maximize use of high-performance computational resources available at leadership-class computing facilities enables these facilities to be leveraged for drug discovery. Here we present the porting, optimization, and validation of the AutoDock-GPU program for the Summit supercomputer, and its application to initial compound screening efforts to target proteins of the SARS-CoV-2 virus responsible for the current COVID-19 pandemic.

LeGrand, Scott↗

Integrated data-driven and experimental approaches to accelerate lead optimization targeting SARS-CoV- 2 main protease

Identification of potential therapeutic candidates can be expedited by integrating computational modeling with domain aware machine learning (ML) approaches followed by experimental validation. Generative deep learning models have been recently developed that can generate thousands of new candidates, but their physiochemical properties are typically not optimized. Using our deep learning models and a scaffold as a starting point, we generated tens of thousands of compounds for SARS-CoV-2 M pro that preserve the core scaffold. Here we utilized and implemented several computational tools such as structural alert and toxicity analysis, high throughput virtual screening, ML-based 3D quantitative structure–activity relationships, multi-parameter optimization, and graph neural networks on libraries of generated candidates to predict biological activity and binding affinity a priori. From these collective computational results, eight promising candidates were identified and tested experimentally using Native Mass Spectrometry (MS) and FRET-based functional assays. Two compounds, with quinazoline-2-thiol and acetylpiperidine core moiety showed IC 50 values in the low micromolar range: 2.95±0.0017 µM and 3.41±0.0015 µM, respectively. The molecular dynamics simulations further highlight that binding of these compounds results in allosteric modulations in the chain B and the interface domains of the M pro . The key fragments from these top hits can be used as input for closed loop lead optimization in the integrated pipeline.

60 APPLIED LIFE SCIENCES↗

An active learning high-throughput microstructure calibration framework for solving inverse structure–process problems in materials informatics

Determining a process–structure–property relationship is the holy grail of materials science, where both computational prediction in the forward direction and materials design in the inverse direction are essential. Problems in materials design are often considered in the context of process–property linkage by bypassing the materials structure, or in the context of structure–property linkage as in microstructure-sensitive design problems. However, there is a lack of research effort in studying materials design problems in the context of process–structure linkage, which has a great implication in reverse engineering. In this paper, given a target microstructure, we propose an active learning high-throughput microstructure calibration framework to derive a set of processing parameters, which can produce an optimal microstructure that is statistically equivalent to the target microstructure. The proposed framework is formulated as a noisy multi-objective optimization problem, where each objective function measures a deterministic or statistical difference of the same microstructure descriptor between a candidate microstructure and a target microstructure. Furthermore, to significantly reduce the physical waiting wall-time, we enable the high-throughput feature of the microstructure calibration framework by adopting an asynchronously parallel Bayesian optimization to exploit high-performance computing resources. Case studies in additive manufacturing and grain growth are used to demonstrate the applicability of the proposed framework, where kinetic Monte Carlo (kMC) simulation is used as a forward predictive model, such that for a given target microstructure, the target processing parameters that produced this microstructure are successfully recovered.

36 MATERIALS SCIENCE↗

EJFAT: Towards Intelligent Compute Destination Load Balancing

To handle increased data flow, Jefferson Lab (JLab) is partnering with ESnet for development of an AI/ML directed compute work Load Balancer (LB) of UDP streamed data. The LB is FPGA based featuring dynamically configurable, low latency and high throughput destination address switching. The LB provides integration of edge and core computing to support JLab experimental programs, the Electron-Ion Collider, as well as data centers of the future. In the ESnet/JLab FPGA Accelerated Transport (EJFAT) initiative, the function of the LB Data Plane (DP) is to redirect data streams to selectable (but unknown to sender) destination hosts based on current worload and within that host to destination ports as a function of sub- stream id. This effects hierarchical scaling, first across compute machines for processing over a series of events and second, across ports so different data source sub-streams may be assigned to different processors for further parallelization. The LB Control Plane (CP) programs the DP using compute farm telemetry to direct and balance workloads across a compute cluster as the operating conditions require. While Proportional/Integrative/Derivative (PID) controllers are often seen in similar applications, here we investigate the feasibility of a Reinforcement Learning (RL) based schedule manager running in the CP to provide dynamic updates to the DP scheduling policy.

Lawrence, David↗