Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “high throughput computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

A modular minicomputer based Navier-Stokes solver

The basic module consists of a minicomputer, low cost peripheral storage device (disk) and a modest number (8-12) of microcomputer modules. A simple arrangement, where the microcomputers are connected to a single time multiplexed bus, only communicating to the host minicomputer, will be efficient. By running the machine in a dedicated mode for long periods of time, it will be possible to obtain a large number of solutions. As such, the device should be useful as a research tool. A scheme is outlined to assemble a number of these computing modules in parallel to decrease computing time. The advantages and disadvantages are discussed of using a number of these systems assembled in a loosely coupled configuration, each independently computing a separate flow, to give a very high throughput.

Steinhoff, J.↗

Speeding up particle track reconstruction using a parallel Kalman filter algorithm

One of the most computationally difficult problems expected for the High-Luminosity Large Hadron Collider (HL-LHC) is determining the trajectory of charged particles during event reconstruction. Algorithms used at the LHC today rely on Kalman filtering, which builds physical trajectories incrementally while incorporating material effects and error estimation. Recognizing the need for faster computational throughput, we have adapted Kalman-filter-based methods for highly parallel, many-core SIMD architectures that are now prevalent in high-performance hardware. In this paper, we discuss the design and performance of the improved tracking algorithm, referred to as mkFit. A key piece of the algorithm is the Matriplex library, containing dedicated code to optimally vectorize operations on small matrices. The physics performance of the mkFit algorithm is comparable to the nominal CMS tracking algorithm when reconstructing tracks from simulated proton-proton collisions within the CMS detector. We study the scaling of the algorithm as a function of the parallel resources utilized and find large speedups both from vectorization and multi-threading. mkFit achieves a speedup of a factor of 6 compared to the nominal algorithm when run in a single-threaded application within the CMS software framework.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Exploring the use of I/O nodes for computation in a MIMD multiprocessor

As parallel systems move into the production scientific-computing world, the emphasis will be on cost-effective solutions that provide high throughput for a mix of applications. Cost effective solutions demand that a system make effective use of all of its resources. Many MIMD multiprocessors today, however, distinguish between 'compute' and 'I/O' nodes, the latter having attached disks and being dedicated to running the file-system server. This static division of responsibilities simplifies system management but does not necessarily lead to the best performance in workloads that need a different balance of computation and I/O. Of course, computational processes sharing a node with a file-system service may receive less CPU time, network bandwidth, and memory bandwidth than they would on a computation-only node. In this paper we begin to examine this issue experimentally. We found that high performance I/O does not necessarily require substantial CPU time, leaving plenty of time for application computation. There were some complex file-system requests, however, which left little CPU time available to the application. (The impact on network and memory bandwidth still needs to be determined.) For applications (or users) that cannot tolerate an occasional interruption, we recommend that they continue to use only compute nodes. For tolerant applications needing more cycles than those provided by the compute nodes, we recommend that they take full advantage of both compute and I/O nodes for computation, and that operating systems should make this possible.

Kotz, David↗

Candidate ferroelectrics via ab initio high-throughput screening of polar materials

Ferroelectrics are a class of polar and switchable functional materials with diverse applications, from microelectronics to energy conversion. Computational searches for new ferroelectric materials have been constrained by accurate prediction of the polarization and switchability with electric field, properties that, in principle, require a comparison with a nonpolar phase whose atomic-scale unit cell is continuously deformable from the polar ground state. For most polar materials, such a higher-symmetry nonpolar phase does not exist or is unknown. Here, we introduce a general high-throughput workflow that screens polar materials as potential ferroelectrics. We demonstrate our workflow on 1978 polar structures in the Materials Project database, for which we automatically generate a nonpolar reference structure using pseudosymmetries, and then compute the polarization difference and energy barrier between polar and nonpolar phases, comparing the predicted values to known ferroelectrics. Focusing on a subset of 182 potential ferroelectrics, we implement a systematic ranking strategy that prioritizes candidates with large polarization and small polar-nonpolar energy differences. To assess stability and synthesizability, we combine information including the computed formation energy above the convex hull, the Inorganic Crystal Structure Database id number, a previously reported machine learning-based synthesizability score, and ab initio phonon band structures. To distinguish between previously reported ferroelectrics, materials known for alternative applications, and lesser-known materials, we combine this ranking with a survey of the existing literature on these candidates through Google Scholar and Scopus databases, revealing ~130 promising materials uninvestigated as ferroelectric. Our workflow and large-scale high-throughput screening lays the groundwork for the discovery of novel ferroelectrics, revealing numerous candidates materials for future experimental and theoretical endeavors.

36 MATERIALS SCIENCE↗

Phase-based velocity extraction method for photonic Doppler velocimetry with potential higher time resolution

We present an extension of the [Takeda et al., J. Opt. Soc. Am. 72, 156 (1982)] phase extraction method to heterodyne photonic Doppler velocimetry applications. The method yields results equivalent to those obtained by the short-time Fourier transform (STFT), while offering potential improvements in time resolution. Unlike STFT, which relies on window functions, such as the Hamming window, that emphasize central data points and diminish the influence of edges, the extended Takeda method utilizes all data uniformly. This uniform treatment allows for the derivation of empirical equations that directly relate velocity error to the actual time resolution rather than to the local analysis duration. The established equation provides a useful metric for both optimizing hardware configuration and guiding data analysis. Simulation and experimental results confirm that, for a given dataset, specifying a target time resolution yields consistent velocity errors for both methods. These findings underscore the Takeda method’s advantages, particularly its potential higher time resolution and reduced computational burden, making it a valuable tool for high-throughput applications such as laser dynamic compression experiments.

Computer simulation↗

A High-Throughput, Adaptive FFT Architecture for FPGA-Based Space-Borne Data Processors

Historically, computationally-intensive data processing for space-borne instruments has heavily relied on ground-based computing resources. But with recent advances in functional densities of Field-Programmable Gate-Arrays (FPGAs), there has been an increasing desire to shift more processing on-board; therefore relaxing the downlink data bandwidth requirements. Fast Fourier Transforms (FFTs) are commonly used building blocks for data processing applications, with a growing need to increase the FFT block size. Many existing FFT architectures have mainly emphasized on low power consumption or resource usage; but as the block size of the FFT grows, the throughput is often compromised first. In addition to power and resource constraints, space-borne digital systems are also limited to a small set of space-qualified memory elements, which typically lag behind the commercially available counterparts in capacity and bandwidth. The bandwidth limitation of the external memory creates a bottleneck for a large, high-throughput FFT design with large block size. In this paper, we present the Multi-Pass Wide Kernel FFT (MPWK-FFT) architecture for a moderately large block size (32K) with considerations to power consumption and resource usage, as well as throughput. We will also show that the architecture can be easily adapted for different FFT block sizes with different throughput and power requirements. The result is completely contained within an FPGA without relying on external memories. Implementation results are summarized.

Nguyen, Kayla↗

Multiscale porosity characterization in additively manufactured polymer nanocomposites using micro-computed tomography

Extrusion-based additive manufacturing (AM) of polymer composites exhibits complex thermally driven phenomena that introduce severe discontinuities in the internal structure across length scales, especially voids or porosity. This study utilizes a high-throughput porosity characterization technique to analyze large datasets from numerous micro-computed tomography (mu CT) scans to capture the influence of AM print parameters on the size, shape, and location of porosity across multiple print layers (up to a few cm) on fused granular fabrication (FGF) printers. The materials investigated include nanocomposite formulations based on commercially relevant nylon-12 and polyether ketone ketone (PEKK) materials comprising nano- or micro- sized fillers. The estimated global porosity follows an inverse linear correlation against the bulk density of the printed samples. Increasing the extrusion multiplier (EM) and the nozzle temperature while decreasing the print speeds reduces the global porosity. Outlier analyses (local porosity morphology) show that faster print speeds and higher extrusion rates result in long, slender inter-layer voids, while lower nozzle temperatures lead to large, symmetrical, inter-bead voids (at the bead junction). Lack of active chamber temperature increases inter-layer and intra-bead voids with a two-fold increase in global porosity. Overall, the micro filler-reinforced composites exhibit higher global porosity than nanofiller-reinforced composites, which is attributed to the increased mismatch in the thermal expansion coefficient between the filler and the polymers used in the study.

36 MATERIALS SCIENCE↗

A flexible and scalable scheme for mixing computed formation energies from different levels of theory

Abstract Computational materials discovery efforts are enabled by large databases of properties derived from high-throughput density functional theory (DFT), which now contain millions of calculations at the generalized gradient approximation (GGA) level of theory. It is now feasible to carry out high-throughput calculations using more accurate methods, such as meta-GGA DFT; however recomputing an entire database with a higher-fidelity method would not effectively leverage the enormous investment of computational resources embodied in existing (GGA) calculations. Instead, we propose here a general procedure by which higher-fidelity, low-coverage calculations (e.g., meta-GGA calculations for selected chemical systems) can be combined with lower-fidelity, high-coverage calculations (e.g., an existing database of GGA calculations) in a robust and scalable manner. We then use legacy PBE(+ U ) GGA calculations and new r 2 SCAN meta-GGA calculations from the Materials Project database to demonstrate that our scheme improves solid and aqueous phase stability predictions, and discuss practical considerations for its implementation.

36 MATERIALS SCIENCE↗

Generic, Extensible, Configurable Push-Pull Framework for Large-Scale Science Missions

The push-pull framework was developed in hopes that an infrastructure would be created that could literally connect to any given remote site, and (given a set of restrictions) download files from that remote site based on those restrictions. The Cataloging and Archiving Service (CAS) has recently been re-architected and re-factored in its canonical services, including file management, workflow management, and resource management. Additionally, a generic CAS Crawling Framework was built based on motivation from Apache s open-source search engine project called Nutch. Nutch is an Apache effort to provide search engine services (akin to Google), including crawling, parsing, content analysis, and indexing. It has produced several stable software releases, and is currently used in production services at companies such as Yahoo, and at NASA's Planetary Data System. The CAS Crawling Framework supports many of the Nutch Crawler's generic services, including metadata extraction, crawling, and ingestion. However, one service that was not ported over from Nutch is a generic protocol layer service that allows the Nutch crawler to obtain content using protocol plug-ins that download content using implementations of remote protocols, such as HTTP, FTP, WinNT file system, HTTPS, etc. Such a generic protocol layer would greatly aid in the CAS Crawling Framework, as the layer would allow the framework to generically obtain content (i.e., data products) from remote sites using protocols such as FTP and others. Augmented with this capability, the Orbiting Carbon Observatory (OCO) and NPP (NPOESS Preparatory Project) Sounder PEATE (Product Evaluation and Analysis Tools Elements) would be provided with an infrastructure to support generic FTP-based pull access to remote data products, obviating the need for any specialized software outside of the context of their existing process control systems. This extensible configurable framework was created in Java, and allows the use of different underlying communication middleware (at present, both XMLRPC, and RMI). In addition, the framework is entirely suitable in a multi-mission environment and is supporting both NPP Sounder PEATE and the OCO Mission. Both systems involve tasks such as high-throughput job processing, terabyte-scale data management, and science computing facilities. NPP Sounder PEATE is already using the push-pull framework to accept hundreds of gigabytes of IASI (infrared atmospheric sounding interferometer) data, and is in preparation to accept CRIMS (Cross-track Infrared Microwave Sounding Suite) data. OCO will leverage the framework to download MODIS, CloudSat, and other ancillary data products for use in the high-performance Level 2 Science Algorithm. The National Cancer Institute is also evaluating the framework for use in sharing and disseminating cancer research data through its Early Detection Research Network (EDRN).

Foster, Brian M.↗

Protein-ligand binding affinity prediction using multi-instance learning with docking structures

Recent advances in 3D structure-based deep learning approaches demonstrate improved accuracy in predicting protein-ligand binding affinity in drug discovery. These methods complement physics-based computational modeling such as molecular docking for virtual high-throughput screening. Despite recent advances and improved predictive performance, most methods in this category primarily rely on utilizing co-crystal complex structures and experimentally measured binding affinities as both input and output data for model training. Nevertheless, co-crystal complex structures are not readily available and the inaccurate predicted structures from molecular docking can degrade the accuracy of the machine learning methods. We introduce a novel structure-based inference method utilizing multiple molecular docking poses for each complex entity. Our proposed method employs multi-instance learning with an attention network to predict binding affinity from a collection of docking poses. We validate our method using multiple datasets, including PDBbind and compounds targeting the main protease of SARS-CoV-2. The results demonstrate that our method leveraging docking poses is competitive with other state-of-the-art inference models that depend on co-crystal structures. This method offers binding affinity prediction without requiring co-crystal structures, thereby increasing its applicability to protein targets lacking such data.

97 MATHEMATICS AND COMPUTING↗

Root system architecture in cereals: progress, challenges and perspective

We report roots are essential multifunctional plant organs involved in water and nutrient uptake, metabolite storage, anchorage, mechanical support, and interaction with the soil environment. Understanding of this ‘hidden half’ provides potential for manipulation of root system architecture (RSA) traits to optimize resource use efficiency and grain yield in cereal crops. Unfortunately, root traits are highly neglected in breeding due to the challenges of phenotyping, but could have large rewards if the variability in RSA traits can be fully exploited. Until now, a plethora of genes have been characterized in detail for their potential role in improving RSA. The use of forward genetics approaches to find sequence variations in genes underpinning desirable RSA would be highly beneficial. Advances in computer vision applications have allowed image-based approaches for high-throughput phenotyping of RSA traits that can be used by any laboratory worldwide to make progress in understanding root function and dissection of the genetics. At the same time, the frontiers of root measurement include non-invasive methods like X-ray computer tomography and magnetic resonance imaging that facilitate new types of temporal studies. Root physiology and ecology are further supported by spatiotemporal root simulation modeling. The discovery of component traits providing improved resilience and yield advantage in target environments is a key necessity for mainstreaming root-based cereal breeding. The integrated use of pan-genome resources, now available in most cereals, coupled with new in-field phenotyping platforms has the potential for precise selection of superior genotypes with improved RSA.

59 BASIC BIOLOGICAL SCIENCES↗

Universal machine learning framework for defect predictions in zinc blende semiconductors

Our article introduces a universal predictive framework for point defect formation energies and charge transition levels in a wide chemical space of zinc blende semiconductors and possible impurity atoms selected from across the periodic table. This framework was developed by leveraging high-throughput quantum mechanical simulations benchmarked using some experimental data from the literature, as well as machine learning (ML)-based regressions techniques that map unique materials descriptors to computed defect properties and yield optimized and generalizable models. Furthermore, the power and utility of these models is revealed through quick predictions for thousands of new defects and screening of low-energy impurities, which may tune the equilibrium conductivity in the semiconductor. This work presents, to our knowledge, the largest density functional theory (DFT) dataset of defect properties in semiconductors and the largest DFT+ML-based screening of point defects in semiconductors to date.

36 MATERIALS SCIENCE↗

Developing machine learning for heterogeneous catalysis with experimental and computational data

Machine learning techniques have emerged as a useful tool for identifying complex patterns and correlations in large datasets, such as associating catalyst performance to its physicochemical properties. In the heterogeneous catalysis communities, machine learning models have mostly been developed using high-throughput quantum chemistry calculations, with only a few case studies resulting in experimentally validated catalyst improvements. This limited success may be due to the use of simplified catalyst structures in computational studies and the lack of comprehensive experimental datasets. In this Review, we bring together studies integrating high-throughput approaches and machine learning for the advancement of solid heterogeneous catalysis, leveraging both experimental and computational data. We systematically analyze trends in the field, based on the descriptors used as model input and output; the materials, devices, or reactions investigated; the dataset size; and the overall achievements. Furthermore, for models reporting unitless R 2 values, we compare the performances based on these mentioned trends.

Computational chemistry↗

Machine learned potential for high-throughput phonon calculations of metal—organic frameworks

Metal–organic frameworks (MOFs) are highly porous and versatile materials studied extensively for applications such as carbon capture and water harvesting. However, computing phonon-mediated properties in MOFs, like thermal expansion and mechanical stability, remains challenging due to the large number of atoms per unit cell, making traditional Density Functional Theory (DFT) methods impractical for high-throughput screening. Recent advances in machine learning potentials have led to foundation atomistic models, such as MACE-MP-0, that accurately predict equilibrium structures but struggle with phonon properties of MOFs. In this work, we developed a workflow for computing phonons in MOFs within the quasi-harmonic approximation with a fine-tuned MACE model, MACE-MP-MOF0. The model was trained on a curated dataset of 127 representative and diverse MOFs. The fine-tuned MACE-MP-MOF0 improves the accuracy of phonon density of states and corrects the imaginary phonon modes of MACE-MP-0, enabling high-throughput phonon calculations with state-of-the-art precision. The model successfully predicts thermal expansion and bulk moduli in agreement with DFT and experimental data for several well-known MOFs. These results highlight the potential of MACE-MP-MOF0 in guiding MOF design for applications in energy storage and thermoelectrics.

Elena, Alin Marin↗

HEPOM: Using Graph Neural Networks for the Accelerated Predictions of Hydrolysis Free Energies in Different pH Conditions

Hydrolysis is a fundamental family of chemical reactions where water facilitates the cleavage of bonds. The process is ubiquitous in biological and chemical systems, owing to water’s remarkable versatility as a solvent. However, accurately predicting the feasibility of hydrolysis through computational techniques is a difficult task, as subtle changes in reactant structure like heteroatom substitutions or neighboring functional groups can influence the reaction outcome. Furthermore, hydrolysis is sensitive to the pH of the aqueous medium, and the same reaction can have different reaction properties at different pH conditions. In this work, we have combined reaction templates and high-throughput ab initio calculations to construct a diverse data set of hydrolysis free energies. The developed framework automatically identifies reaction centers, generates hydrolysis products, and utilizes a trained graph neural network (GNN) model to predict ΔG values for all potential hydrolysis reactions in a given molecule. The long-term goal of the work is to develop a data-driven, computational tool for high-throughput screening of pH-specific hydrolytic stability and the rapid prediction of reaction products, which can then be applied in a wide array of applications including chemical recycling of polymers and ion-conducting membranes for clean energy generation and storage.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Generic High Bandwidth Data Acquisition Card for Physics Experiments

In high energy physics and nuclear physics experiments particularly the ones based on particle accelerator, the data rate from the detector is usually in the order of Terabytes per second. This high throughput data from detector front-end electronics need to be transmitted to the back-end computing farm for high level event selection and building. A Data Acquisition (DAQ) system with features of high-density, scalable, easily upgradeable is crucial to simplify the readout architecture of whole experiment. This paper will introduce the design of a generic high bandwidth PCIe card which can be used as the important input output card in a scalable DAQ system. It can factorize front-end electronics from data handling, and reduce amount of custom hardware in favor of scalable detectorindependent commercial hardware and software. Besides the 48 channels of bidirectional high speed fiber optical links with frontends, it also supports to synchronize with the experiment timing system, and to fanout the clock and trigger information with a fixed latency to the front-end electronics.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Spectral interferometry-based microwave-frequency vibrometry for integrated acoustic wave devices

Microwave phononics is a promising platform for sensing, computing, and quantum information science; thus, sensitive and high-throughput characterization tools are needed not only for device verification and optimization but also for revealing transient and nonlinear dynamics. Existing interferometric optical vibrometers for 2D mapping are challenged by operating point stabilization, surface reflectivity contrast, and long acquisition time. Here, we use spectral interferometry, which is insensitive to these factors and utilizes a continuous raster scanning scheme for vibration mapping with high throughput. We intensity-modulate our broadband light source with an electro-optic modulator to resolve vibrations at microwave frequencies. Our system requires no fast photodetector or digitizer operating in the microwave frequency range. We image the 1 GHz vibration field of a 300 × 150 µm 2 area of an entire surface acoustic wave device in 10 min with simultaneous surface profilometry. Our system has a vibration sensitivity of 120 fm/sqrt(Hz) and a linear throughput of 0.77 mm/s on the chip surface. The technique offers capabilities for characterizing a wide range of acoustic wave and micromechanical devices to better understand their behavior and performance.

47 OTHER INSTRUMENTATION↗