Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “high throughput computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Machine Learning Thermodynamics And Kinetics of Defects For Accelerated Materials Discovery

Atomistic defects play a pivotal role in functional and structural materials’ performance across a myriad of technology applications. Quantitative prediction of the thermodynamics and kinetics of defect formation and migration, respectively, typically requires accurate but expensive first-principles approaches, such as density functional theory (DFT). Their computational expense limits the throughput needed to perform high-throughput materials discovery/screening exercises or to perform materials modeling tasks relying on extensive sampling techniques. Therefore, in this Sandia National Laboratories Laboratory Directed Research and Development (LDRD) project (Project #229366), we developed a variety of machine learning techniques, trained on density functional theory calculations, to accelerate the discovery and modeling of materials in which vacancy and interstitial defects primarily dictate material performance. These include applications such as metal oxides for water-splitting or mixed ionic-electronic conduction, metal hydrides for hydrogen storage, and transition metal dichalcogenides for electronics, and the approaches developed herein can further be applied to many other domains that similarly depend on materials’ thermodynamic and kinetic defect properties for their desired functionality.

36 MATERIALS SCIENCE↗

Optimal Decision Making in High-Throughput Virtual Screening Pipelines

ABSTRACT Effective selection of the potential candidates that meet certain conditions in a tremendously large search space has been one of the major concerns in many real-world applications. In addition to the nearly infinitely large search space, rigorous evaluation of a sample based on the reliable experimental or computational platform is often prohibitively expensive, making the screening problem more challenging. In such a case, constructing a high-throughput screening (HTS) pipeline that pre-sifts the samples expected to be potential candidates through the efficient earlier stages, results in a significant amount of savings in resources. However, to the best of our knowledge, despite many successful applications, no one has studied optimal pipeline design or optimal pipeline operations. In this study, we propose two optimization frameworks, applying to most (if not all) screening campaigns involving experimental or/and computational evaluations, for optimally determining the screening thresholds of an HTS pipeline. We validate the proposed frameworks on both analytic and practical scenarios. In particular, we consider the optimal computational campaign for the long non-coding RNA (lncRNA) classification as a practical example. To accomplish this, we built the high-throughput virtual screening (HTVS) pipeline for classifying the lncRNA. The simulation results demonstrate that the proposed frameworks significantly reduce the effective selection cost per potential candidate and make the HTS pipelines less sensitive to their structural variations. In addition to the validation, we provide insights on constructing a better HTS pipeline based on the simulation results.

97 MATHEMATICS AND COMPUTING↗

Speeding up particle track reconstruction using a parallel Kalman filter algorithm

One of the most computationally difficult problems expected for the High-Luminosity Large Hadron Collider (HL-LHC) is determining the trajectory of charged particles during event reconstruction. Algorithms used at the LHC today rely on Kalman filtering, which builds physical trajectories incrementally while incorporating material effects and error estimation. Recognizing the need for faster computational throughput, we have adapted Kalman-filter-based methods for highly parallel, many-core SIMD architectures that are now prevalent in high-performance hardware. In this paper, we discuss the design and performance of the improved tracking algorithm, referred to as mkFit. A key piece of the algorithm is the Matriplex library, containing dedicated code to optimally vectorize operations on small matrices. The physics performance of the mkFit algorithm is comparable to the nominal CMS tracking algorithm when reconstructing tracks from simulated proton-proton collisions within the CMS detector. We study the scaling of the algorithm as a function of the parallel resources utilized and find large speedups both from vectorization and multi-threading. mkFit achieves a speedup of a factor of 6 compared to the nominal algorithm when run in a single-threaded application within the CMS software framework.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Candidate ferroelectrics via ab initio high-throughput screening of polar materials

Ferroelectrics are a class of polar and switchable functional materials with diverse applications, from microelectronics to energy conversion. Computational searches for new ferroelectric materials have been constrained by accurate prediction of the polarization and switchability with electric field, properties that, in principle, require a comparison with a nonpolar phase whose atomic-scale unit cell is continuously deformable from the polar ground state. For most polar materials, such a higher-symmetry nonpolar phase does not exist or is unknown. Here, we introduce a general high-throughput workflow that screens polar materials as potential ferroelectrics. We demonstrate our workflow on 1978 polar structures in the Materials Project database, for which we automatically generate a nonpolar reference structure using pseudosymmetries, and then compute the polarization difference and energy barrier between polar and nonpolar phases, comparing the predicted values to known ferroelectrics. Focusing on a subset of 182 potential ferroelectrics, we implement a systematic ranking strategy that prioritizes candidates with large polarization and small polar-nonpolar energy differences. To assess stability and synthesizability, we combine information including the computed formation energy above the convex hull, the Inorganic Crystal Structure Database id number, a previously reported machine learning-based synthesizability score, and ab initio phonon band structures. To distinguish between previously reported ferroelectrics, materials known for alternative applications, and lesser-known materials, we combine this ranking with a survey of the existing literature on these candidates through Google Scholar and Scopus databases, revealing ~130 promising materials uninvestigated as ferroelectric. Our workflow and large-scale high-throughput screening lays the groundwork for the discovery of novel ferroelectrics, revealing numerous candidates materials for future experimental and theoretical endeavors.

36 MATERIALS SCIENCE↗

Phase-based velocity extraction method for photonic Doppler velocimetry with potential higher time resolution

We present an extension of the [Takeda et al., J. Opt. Soc. Am. 72, 156 (1982)] phase extraction method to heterodyne photonic Doppler velocimetry applications. The method yields results equivalent to those obtained by the short-time Fourier transform (STFT), while offering potential improvements in time resolution. Unlike STFT, which relies on window functions, such as the Hamming window, that emphasize central data points and diminish the influence of edges, the extended Takeda method utilizes all data uniformly. This uniform treatment allows for the derivation of empirical equations that directly relate velocity error to the actual time resolution rather than to the local analysis duration. The established equation provides a useful metric for both optimizing hardware configuration and guiding data analysis. Simulation and experimental results confirm that, for a given dataset, specifying a target time resolution yields consistent velocity errors for both methods. These findings underscore the Takeda method’s advantages, particularly its potential higher time resolution and reduced computational burden, making it a valuable tool for high-throughput applications such as laser dynamic compression experiments.

Computer simulation↗

Multiscale porosity characterization in additively manufactured polymer nanocomposites using micro-computed tomography

Extrusion-based additive manufacturing (AM) of polymer composites exhibits complex thermally driven phenomena that introduce severe discontinuities in the internal structure across length scales, especially voids or porosity. This study utilizes a high-throughput porosity characterization technique to analyze large datasets from numerous micro-computed tomography (mu CT) scans to capture the influence of AM print parameters on the size, shape, and location of porosity across multiple print layers (up to a few cm) on fused granular fabrication (FGF) printers. The materials investigated include nanocomposite formulations based on commercially relevant nylon-12 and polyether ketone ketone (PEKK) materials comprising nano- or micro- sized fillers. The estimated global porosity follows an inverse linear correlation against the bulk density of the printed samples. Increasing the extrusion multiplier (EM) and the nozzle temperature while decreasing the print speeds reduces the global porosity. Outlier analyses (local porosity morphology) show that faster print speeds and higher extrusion rates result in long, slender inter-layer voids, while lower nozzle temperatures lead to large, symmetrical, inter-bead voids (at the bead junction). Lack of active chamber temperature increases inter-layer and intra-bead voids with a two-fold increase in global porosity. Overall, the micro filler-reinforced composites exhibit higher global porosity than nanofiller-reinforced composites, which is attributed to the increased mismatch in the thermal expansion coefficient between the filler and the polymers used in the study.

36 MATERIALS SCIENCE↗

A flexible and scalable scheme for mixing computed formation energies from different levels of theory

Abstract Computational materials discovery efforts are enabled by large databases of properties derived from high-throughput density functional theory (DFT), which now contain millions of calculations at the generalized gradient approximation (GGA) level of theory. It is now feasible to carry out high-throughput calculations using more accurate methods, such as meta-GGA DFT; however recomputing an entire database with a higher-fidelity method would not effectively leverage the enormous investment of computational resources embodied in existing (GGA) calculations. Instead, we propose here a general procedure by which higher-fidelity, low-coverage calculations (e.g., meta-GGA calculations for selected chemical systems) can be combined with lower-fidelity, high-coverage calculations (e.g., an existing database of GGA calculations) in a robust and scalable manner. We then use legacy PBE(+ U ) GGA calculations and new r 2 SCAN meta-GGA calculations from the Materials Project database to demonstrate that our scheme improves solid and aqueous phase stability predictions, and discuss practical considerations for its implementation.

36 MATERIALS SCIENCE↗

Protein-ligand binding affinity prediction using multi-instance learning with docking structures

Recent advances in 3D structure-based deep learning approaches demonstrate improved accuracy in predicting protein-ligand binding affinity in drug discovery. These methods complement physics-based computational modeling such as molecular docking for virtual high-throughput screening. Despite recent advances and improved predictive performance, most methods in this category primarily rely on utilizing co-crystal complex structures and experimentally measured binding affinities as both input and output data for model training. Nevertheless, co-crystal complex structures are not readily available and the inaccurate predicted structures from molecular docking can degrade the accuracy of the machine learning methods. We introduce a novel structure-based inference method utilizing multiple molecular docking poses for each complex entity. Our proposed method employs multi-instance learning with an attention network to predict binding affinity from a collection of docking poses. We validate our method using multiple datasets, including PDBbind and compounds targeting the main protease of SARS-CoV-2. The results demonstrate that our method leveraging docking poses is competitive with other state-of-the-art inference models that depend on co-crystal structures. This method offers binding affinity prediction without requiring co-crystal structures, thereby increasing its applicability to protein targets lacking such data.

97 MATHEMATICS AND COMPUTING↗

Root system architecture in cereals: progress, challenges and perspective

We report roots are essential multifunctional plant organs involved in water and nutrient uptake, metabolite storage, anchorage, mechanical support, and interaction with the soil environment. Understanding of this ‘hidden half’ provides potential for manipulation of root system architecture (RSA) traits to optimize resource use efficiency and grain yield in cereal crops. Unfortunately, root traits are highly neglected in breeding due to the challenges of phenotyping, but could have large rewards if the variability in RSA traits can be fully exploited. Until now, a plethora of genes have been characterized in detail for their potential role in improving RSA. The use of forward genetics approaches to find sequence variations in genes underpinning desirable RSA would be highly beneficial. Advances in computer vision applications have allowed image-based approaches for high-throughput phenotyping of RSA traits that can be used by any laboratory worldwide to make progress in understanding root function and dissection of the genetics. At the same time, the frontiers of root measurement include non-invasive methods like X-ray computer tomography and magnetic resonance imaging that facilitate new types of temporal studies. Root physiology and ecology are further supported by spatiotemporal root simulation modeling. The discovery of component traits providing improved resilience and yield advantage in target environments is a key necessity for mainstreaming root-based cereal breeding. The integrated use of pan-genome resources, now available in most cereals, coupled with new in-field phenotyping platforms has the potential for precise selection of superior genotypes with improved RSA.

59 BASIC BIOLOGICAL SCIENCES↗

Universal machine learning framework for defect predictions in zinc blende semiconductors

Our article introduces a universal predictive framework for point defect formation energies and charge transition levels in a wide chemical space of zinc blende semiconductors and possible impurity atoms selected from across the periodic table. This framework was developed by leveraging high-throughput quantum mechanical simulations benchmarked using some experimental data from the literature, as well as machine learning (ML)-based regressions techniques that map unique materials descriptors to computed defect properties and yield optimized and generalizable models. Furthermore, the power and utility of these models is revealed through quick predictions for thousands of new defects and screening of low-energy impurities, which may tune the equilibrium conductivity in the semiconductor. This work presents, to our knowledge, the largest density functional theory (DFT) dataset of defect properties in semiconductors and the largest DFT+ML-based screening of point defects in semiconductors to date.

36 MATERIALS SCIENCE↗

Developing machine learning for heterogeneous catalysis with experimental and computational data

Machine learning techniques have emerged as a useful tool for identifying complex patterns and correlations in large datasets, such as associating catalyst performance to its physicochemical properties. In the heterogeneous catalysis communities, machine learning models have mostly been developed using high-throughput quantum chemistry calculations, with only a few case studies resulting in experimentally validated catalyst improvements. This limited success may be due to the use of simplified catalyst structures in computational studies and the lack of comprehensive experimental datasets. In this Review, we bring together studies integrating high-throughput approaches and machine learning for the advancement of solid heterogeneous catalysis, leveraging both experimental and computational data. We systematically analyze trends in the field, based on the descriptors used as model input and output; the materials, devices, or reactions investigated; the dataset size; and the overall achievements. Furthermore, for models reporting unitless R 2 values, we compare the performances based on these mentioned trends.

Computational chemistry↗

Machine learned potential for high-throughput phonon calculations of metal—organic frameworks

Metal–organic frameworks (MOFs) are highly porous and versatile materials studied extensively for applications such as carbon capture and water harvesting. However, computing phonon-mediated properties in MOFs, like thermal expansion and mechanical stability, remains challenging due to the large number of atoms per unit cell, making traditional Density Functional Theory (DFT) methods impractical for high-throughput screening. Recent advances in machine learning potentials have led to foundation atomistic models, such as MACE-MP-0, that accurately predict equilibrium structures but struggle with phonon properties of MOFs. In this work, we developed a workflow for computing phonons in MOFs within the quasi-harmonic approximation with a fine-tuned MACE model, MACE-MP-MOF0. The model was trained on a curated dataset of 127 representative and diverse MOFs. The fine-tuned MACE-MP-MOF0 improves the accuracy of phonon density of states and corrects the imaginary phonon modes of MACE-MP-0, enabling high-throughput phonon calculations with state-of-the-art precision. The model successfully predicts thermal expansion and bulk moduli in agreement with DFT and experimental data for several well-known MOFs. These results highlight the potential of MACE-MP-MOF0 in guiding MOF design for applications in energy storage and thermoelectrics.

Elena, Alin Marin↗

HEPOM: Using Graph Neural Networks for the Accelerated Predictions of Hydrolysis Free Energies in Different pH Conditions

Hydrolysis is a fundamental family of chemical reactions where water facilitates the cleavage of bonds. The process is ubiquitous in biological and chemical systems, owing to water’s remarkable versatility as a solvent. However, accurately predicting the feasibility of hydrolysis through computational techniques is a difficult task, as subtle changes in reactant structure like heteroatom substitutions or neighboring functional groups can influence the reaction outcome. Furthermore, hydrolysis is sensitive to the pH of the aqueous medium, and the same reaction can have different reaction properties at different pH conditions. In this work, we have combined reaction templates and high-throughput ab initio calculations to construct a diverse data set of hydrolysis free energies. The developed framework automatically identifies reaction centers, generates hydrolysis products, and utilizes a trained graph neural network (GNN) model to predict ΔG values for all potential hydrolysis reactions in a given molecule. The long-term goal of the work is to develop a data-driven, computational tool for high-throughput screening of pH-specific hydrolytic stability and the rapid prediction of reaction products, which can then be applied in a wide array of applications including chemical recycling of polymers and ion-conducting membranes for clean energy generation and storage.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Generic High Bandwidth Data Acquisition Card for Physics Experiments

In high energy physics and nuclear physics experiments particularly the ones based on particle accelerator, the data rate from the detector is usually in the order of Terabytes per second. This high throughput data from detector front-end electronics need to be transmitted to the back-end computing farm for high level event selection and building. A Data Acquisition (DAQ) system with features of high-density, scalable, easily upgradeable is crucial to simplify the readout architecture of whole experiment. This paper will introduce the design of a generic high bandwidth PCIe card which can be used as the important input output card in a scalable DAQ system. It can factorize front-end electronics from data handling, and reduce amount of custom hardware in favor of scalable detectorindependent commercial hardware and software. Besides the 48 channels of bidirectional high speed fiber optical links with frontends, it also supports to synchronize with the experiment timing system, and to fanout the clock and trigger information with a fixed latency to the front-end electronics.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Spectral interferometry-based microwave-frequency vibrometry for integrated acoustic wave devices

Microwave phononics is a promising platform for sensing, computing, and quantum information science; thus, sensitive and high-throughput characterization tools are needed not only for device verification and optimization but also for revealing transient and nonlinear dynamics. Existing interferometric optical vibrometers for 2D mapping are challenged by operating point stabilization, surface reflectivity contrast, and long acquisition time. Here, we use spectral interferometry, which is insensitive to these factors and utilizes a continuous raster scanning scheme for vibration mapping with high throughput. We intensity-modulate our broadband light source with an electro-optic modulator to resolve vibrations at microwave frequencies. Our system requires no fast photodetector or digitizer operating in the microwave frequency range. We image the 1 GHz vibration field of a 300 × 150 µm 2 area of an entire surface acoustic wave device in 10 min with simultaneous surface profilometry. Our system has a vibration sensitivity of 120 fm/sqrt(Hz) and a linear throughput of 0.77 mm/s on the chip surface. The technique offers capabilities for characterizing a wide range of acoustic wave and micromechanical devices to better understand their behavior and performance.

47 OTHER INSTRUMENTATION↗

Response of Subsurface Nitrogen-Cycling Microbial Communities to Environmental Fluctuations (Final Technical Report)

Riparian floodplains are dynamic ecosystems linking terrestrial and riverine systems. These floodplains experience hydrological shifts such as changes in water table height, flooding, and drought and can be ‘hotspots’ of biogeochemical cycling due to shifting sediment moisture (and saturation) and subsurface exchanges of water, nutrients, and other compounds across different sediment layers. Subsurface microbial communities are the primary drivers of biogeochemical processes in floodplains, and thus their structure and function can directly influence both surface and groundwater quality. The microbial nitrogen (N) cycle is particularly important in floodplains as it affects nutrient availability and removal. Two functional guilds of chemoautotrophic (i.e. CO2-fixing) microorganisms are responsible for the first oxidative step of the N cycle, nitrification: ammonia-oxidizing archaea (AOA) and bacteria (AOB) catalyze the oxidation of ammonia to nitrite, while nitrite-oxidizing bacteria (NOB) oxidize nitrite to nitrate. Despite the critical role nitrification plays in N-cycling in both terrestrial and aquatic ecosystems, our understanding of the diversity, ecophysiology, and activity of nitrifying organisms in subsurface floodplain soils/sediments is extremely limited. To help address this critical knowledge gap, the overarching goal of this project was to determine how shifts in key environmental parameters and gradients impact microbial N-cycling communities/processes, with particular emphasis on nitrification, within hydrologically-variable floodplain sediments in the Wind River Basin near Riverton, Wyoming. The three specific objectives of this project were to: (1) to associate in situ environmental drivers of N cycling with distinct functional guilds; (2) determine the guild response to variation in key ecosystem drivers; and (3) develop a dynamic ecosystem model of the microbial N cycle with the Riverton subsurface using community genomic and biogeochemical data collected in the first two objectives. Over the course of this project, we employed both 16S rRNA gene amplicon sequencing and genome-resolved metagenomics to examine the phylogenetic diversity and metabolic potential of subsurface nitrifier communities within 68 samples collected across multiple sites, depths, and time points within the Riverton floodplain, allowing for both spatial and temporal investigations at different scales. This project benefitted tremendously from recent advances in high-throughput sequencing technologies coupled with dramatic improvements in the computational tools and algorithms available for analyzing such large, complex genomic datasets. By pairing these cutting-edge genomic approaches with depth-resolved sampling and detailed geochemical analyses of the Riverton floodplain, we have gained novel insights into the structure and function of subsurface nitrifier communities in relation to both hydrology and biogeochemistry. This project resulted in the most detailed and comprehensive characterization of N-cycling floodplain microbial communities to date and will hopefully inspire and pave the way for future studies using similar approaches in other floodplains. Indeed, such information is critical for understanding subsurface biogeochemical cycling and how elemental stores are altered from perturbations initiated by the water cycle within floodplains. Finally, because of the terrestrial-aquatic nature of the Riverton floodplain, results from this project are also of relevance to disciplines such as soil science, estuarine science, limnology & oceanography, biogeochemistry, geobiology, environmental engineering, as well as genomics and data science.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning-Guided Identification of PET Hydrolases from Natural Diversity

The enzymatic depolymerization of poly(ethylene terephthalate) (PET) is emerging as a leading chemical recycling technology for waste polyester. As part of this endeavor, new candidate enzymes identified from natural diversity can serve as useful starting points for enzyme evolution and engineering. In this study, we improved upon HMM searches by applying an iterative machine learning strategy to identify 400 putative PET-degrading enzymes (PET hydrolases) from naturally occurring homologs. Using high-throughput (HTP) experimental techniques, we successfully expressed and purified >200 enzyme candidates and assayed them for PET hydrolysis activity as a function of pH, temperature, and substrate crystallinity. From this library, we discovered 91 previously unknown PET hydrolases, 35 of which retain activity at pH 4.5 on crystalline material, which are conditions relevant to developing more efficient commercial processes. Notably, four enzymes showed equal to or higher activity than LCC-ICCG, a benchmark PET hydrolase, at this challenging condition in our screening assay, and 11 of which have pH optima <7. Using these data, we identified regions of PETases statistically correlated to activity at lower pH. We additionally investigated the effect of condition-specific activity data on trained machine learning predictors and found a precision (putative hit rate) improvement of up to 30% compared to a Hidden Markov Model alone. Our findings show that by pointing enzyme discovery toward conditions of interest with multiple rounds of experimental and machine learning, we can discover large sets of active enzymes and explore factors associated with activity at those conditions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗