Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “streaming algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Transonic Airfoil Analysis

Program uses fast iteration scheme for solving transonic flow field around arbitrary airfoils. Transonic Airfoil Analysis Computer Code, TAIR, employs fast, fully implicit algorithm to solve conservative full-potiential equation for steady transonic flow field about arbitrary airfoil immersed in subsonic free stream. TAIR written in FORTRAN IV.

Holst, T. L.↗

Intelligent Experiments through Real-Time AI: Fast Data Processing and Autonomous Detector Control for High-Energy Nuclear Experiments

The aim of this project is to develop software and hardware for fast real-time data processing and autonomous detector control and calibration for the sPHENIX and the future EIC experiments. Below summarizes Georgia Tech team efforts in the past year: 1. We developed a real-time clustering algorithm and FPGA-based pipeline architecture for processing fired pixel data from ALPIDE sensors in sPHENIX experiments. Our Columnar Clustering Co-Design introduces a hardware-aware, stream-friendly approach that segments pixel data by column pairs using a Column Pair Clustering (CPC) strategy, followed by Cluster Stitching to merge adjacent subclusters. Implemented in Vitis HLS, the pipeline comprises five stages—read-in, subclustering, stitching, analysis, and write-out—connected by tagged HLS streams with custom end-of-event signaling for robust synchronization. We designed a pipelined dataflow model optimized for throughput, low latency, and minimal buffering, enabling scalable clustering across events of arbitrary size. Our system maintains spatial precision via center-of-mass and shape key extraction and efficiently handles edge cases such as fragmented or nested clusters. Compared against DBSCAN in both software and hardware, our approach demonstrates competitive performance under FPGA constraints. 2. We also conducted a comprehensive algorithm-to-hardware co-design of connected component analysis tailored for sPHENIX experiments, focusing on real-time, low-latency processing using FPGAs and High-Level Synthesis (HLS). Starting from a Python-based particle tracking pipeline, the team translated the core logic—graph traversal via DFS and Union-Find—into an HLS-compatible C++ model, replacing dynamic memory and recursion with static arrays and pipelined control flow. The final design includes a fully streamed and dataflow-compatible Union-Find kernel optimized across five iterations, incorporating loop pipelining, array partitioning, AXI/FIFO interface tuning, and function flattening. Experimental results show up to 14.8× speedup over the CPU baseline, reducing per-graph latency to 1.58 μs and demonstrating strong resource efficiency with only ~7k LUTs and zero BRAM usage. The design maintains functional correctness against the Python reference using a Python-based C-simulation framework and Mean Squared Error metrics. This work validates the potential of HLS-driven FPGA designs for edge-level HEP data acquisition, laying a scalable foundation for future integration with real-time detector pipelines and multi-graph processing systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Single-drop reactive extraction/extractive reaction with forced convective diffusion and interphase mass transfer

An algorithm has been developed for the forced convective diffusion-reaction problem for convection inside and outside a droplet by a recirculating flow field hydrodynamically coupled at the droplet interface with an external flow field that at infinity becomes a uniform streaming flow. The concentration field inside the droplet is likewise coupled with that outside by boundary conditions at the interface. A chemical reaction can take place either inside or outside the droplet or reactions can take place in both phases. The algorithm has been implemented and results are shown here for the case of no reaction and for the case of an external first order reaction, both for unsteady behavior. For pure interphase mass transfer, concentration isocontours, local and average Sherwood numbers, and average droplet concentrations have been obtained as a function of the physical properties and external flow field. For mass transfer enhanced by an external reaction, in addition to the above forms of results, we present the enhancement factor, with the results now also depending upon the (dimensionless) rate of reaction.

Kleinman, Leonid S.↗

On-Sensor Data Filtering using Neuromorphic Computing for High Energy Physics Experiments

This work describes the investigation of neuromorphic computing-based spiking neural network (SNN) models used to filter data from sensor electronics in high energy physics experiments conducted at the High Luminosity Large Hadron Collider. We present our approach for developing a compact neuromorphic model that filters out the sensor data based on the particle's transverse momentum with the goal of reducing the amount of data being sent to the downstream electronics. The incoming charge waveforms are converted to streams of binary-valued events, which are then processed by the SNN. We present our insights on the various system design choices - from data encoding to optimal hyperparameters of the training algorithm - for an accurate and compact SNN optimized for hardware deployment. Our results show that an SNN trained with an evolutionary algorithm and an optimized set of hyperparameters obtains a signal efficiency of about 91% with nearly half as many parameters as a deep neural network.

R. Kulkarni, Shruti↗

The status and prospects of materials for carbon capture technologies

In order to combat climate change, carbon dioxide (CO 2 ) emissions from industry, transportation, buildings, and other sources need to be captured and long-term stored. Decarbonization of these sources requires special types of materials that have high affinities for CO 2 . Potassium hydroxide is a benchmark aqueous sorbent that reacts with CO 2 to convert it into K 2 CO 3 and subsequently precipitated as CaCO 3 . Another class of carbon capture materials is solid sorbents that are usually functionalized with amines or have natural affinities for CO 2 . The next wave of materials for carbon capture under investigation includes activated carbon, metal–organic frameworks, zeolites, carbon nanotubes, and ionic liquids. In this issue of MRS Bulletin, some of these materials are highlighted, including solvents and sorbents, membranes, ionic liquids, and hydrides. Other materials that can capture CO2 from low concentrations of gas streams, such as air (direct air capture) are also discussed. Also covered in this issue are machine learning-based computer algorithms developed with the goal to speed up the progress of carbon capture materials development, and to design advanced materials with high CO 2 capacity, improved capture and release kinetics, and improved cyclic durability.

36 MATERIALS SCIENCE↗

Event Definition for the Automated Detection of Nuclear Proliferation Activities

In FY2020, Savannah River National Laboratory (SRNL) in collaboration with the Discovery Analytics Center (DAC) at Virginia Polytechnic Institute and State University (VT) began developing a demonstration prototype system that uses multiple machine learning and data analytic methods on largescale open data sources to identify new, developing, or undeclared nuclear programs. One of the most challenging aspects of applying machine learning techniques to such a problem is the high likelihood of extremely sparse data from disparate sources. To overcome this challenge, the current work will use a strategic combination of supervised, semi-supervised, and unsupervised learning techniques to ingest and fuse data streams to make a forecast of nuclear activities in a targeted geospatial location. Identifying potential data sources and training supervised learning algorithms is dependent upon the development of a robust foundation of targeted event domains that fundamentally define the nuclear activities of interest. This report documents the definition of a hierarchical structure for both nuclear activity and event domains that will be used to guide the research team in development or use of existing semantic dictionaries that are instrumental to searching, parsing, and categorizing events for the forecasting system’s use.

97 MATHEMATICS AND COMPUTING↗

Computation of the inviscid supersonic flow about cones at large angles of attack by a floating discontinuity approach

The technique of floating shock fitting is adapted to the computation of the inviscid flowfield about circular cones in a supersonic free stream at angles of attack that exceed the cone half-angle. The resulting equations are applicable over the complete range of free-stream Mach numbers, angles of attack and cone half-angles for which the bow shock is attached. A finite difference algorithm is used to obtain the solution by an unsteady relaxation approach. The bow shock, embedded cross-flow shock, and vortical singularity in the leeward symmetry plane are treated as floating discontinuities in a fixed computational mesh. Where possible, the flowfield is partitioned into windward, shoulder, and leeward regions with each region computed separately to achieve maximum computational efficiency. An alternative shock fitting technique which treats the bow shock as a computational boundary is developed and compared with the floating-fitting approach. Several surface boundary condition schemes are also analyzed.

Daywitt, J.↗

A computational investigation of supersonic axisymmetric flow over boattails containing a centered propulsive jet

The influence of underexpanded jets on a supersonic afterbody flow field is investigated using computational techniques. The thin-shear-layer formulation of the compressible, Reynolds-averaged Navier-Stokes equations is solved using a time-dependent, implicit numerical algorithm. Solutions are obtained for supersonic flow over an axisymmetric conical afterbody containing a centered propulsive jet where the free-stream Mach number is 2.0 and the jet exit Mach number is 2.5. Exhaust-jet static pressures are considered in the range of 2 to 9 times the free-stream static pressure and with nozzle-exit half-angles from 15 deg to 43 deg. Comparisons are made with experimental results for base pressure, separation distance, afterbody pressure distribution, anf flow-field structure. Although good quantitative agreement with experimental separation distance and base pressure level is not observed, the parametric trends induced by exhaust-jet pressure level and nozzle-exit angle are well predicted, as are the flow-field details in the vicinity of the afterbody and in the exhaust plume.

Deiwert, G. S.↗

Iterative solution of large, sparse linear systems on a static data flow architecture - Performance studies

The applicability of static data flow architectures to the iterative solution of sparse linear systems of equations is investigated. An analytic performance model of a static data flow computation is developed. This model includes both spatial parallelism, concurrent execution in multiple PE's, and pipelining, the streaming of data from array memories through the PE's. The performance model is used to analyze a row partitioned iterative algorithm for solving sparse linear systems of algebraic equations. Based on this analysis, design parameters for the static data flow architecture as a function of matrix sparsity and dimension are proposed.

Reed, D. A.↗

Computational fluid mechanics

Two papers are included in this progress report. In the first, the compressible Navier-Stokes equations have been used to compute leading edge receptivity of boundary layers over parabolic cylinders. Natural receptivity at the leading edge was simulated and Tollmien-Schlichting waves were observed to develop in response to an acoustic disturbance, applied through the farfield boundary conditions. To facilitate comparison with previous work, all computations were carried out at a free stream Mach number of 0.3. The spatial and temporal behavior of the flowfields are calculated through the use of finite volume algorithms and Runge-Kutta integration. The results are dominated by strong decay of the Tollmien-Schlichting wave due to the presence of the mean flow favorable pressure gradient. The effects of numerical dissipation, forcing frequency, and nose radius are studied. The Strouhal number is shown to have the greatest effect on the unsteady results. In the second paper, a transition model for low-speed flows, previously developed by Young et al., which incorporates first-mode (Tollmien-Schlichting) disturbance information from linear stability theory has been extended to high-speed flow by incorporating the effects of second mode disturbances. The transition model is incorporated into a Reynolds-averaged Navier-Stokes solver with a one-equation turbulence model. Results using a variable turbulent Prandtl number approach demonstrate that the current model accurately reproduces available experimental data for first and second-mode dominated transitional flows. The performance of the present model shows significant improvement over previous transition modeling attempts.

Hassan, H. A.↗

Single-drop reactive extraction/extractive reaction with forced convective diffusion and interphase mass transfer

An algorithm has been developed for time-dependent forced convective diffusion-reaction having convection by a recirculating flow field within the drop that is hydrodynamically coupled at the interface with a convective external flow field that at infinity becomes a uniform free-streaming flow. The concentration field inside the droplet is likewise coupled with that outside by boundary conditions at the interface. A chemical reaction can take place either inside or outside the droplet, or reactions can take place in both phases. The algorithm has been implemented, and for comparison results are shown here for the case of no reaction in either phase and for the case of an external first order reaction, both for unsteady behavior. For pure interphase mass transfer, concentration isocontours, local and average Sherwood numbers, and average droplet concentrations have been obtained as a function of the physical properties and external flow field. For mass transfer enhanced by an external reaction, in addition to the above forms of results, we present the enhancement factor, with the results now also depending upon the (dimensionless) rate of reaction.

Kleinman, Leonid S.↗

Contaminant Investigation and Pre‐Processing Opportunities for Textile‐To‐Textile Recycling

Millions of metric tons of textiles are landfilled or incinerated each year in the United States, with less than 1% of textiles recycled into new clothing or fabrics. To counter this trend, a growing number of companies and researchers are exploring how a circular economy can be applied to support textile‐to‐textile recycling. A significant barrier they face comes down to quickly and efficiently extracting pure feedstock material from post‐consumer garments that feature a mix of natural and synthetic fibers. Textile recyclers prefer pure feedstocks, as working with mixed sources typically means lower throughput, higher risk of equipment failure, and diminished business margins. To facilitate a circular economy for textiles, methods, and technologies are needed that can efficiently separate out materials and contaminants from end‐of‐life textiles to increase the flow of pure feedstocks to recyclers. This paper summarizes findings from interviews with a cross section of textile recyclers and from a review of literature to define basic feedstock requirements. In addition to our qualitative research, we deconstruct a bale of post‐consumer textiles and analyze them using computer‐vision imaging, Fourier transform infrared spectroscopy (FTIR), and machine learning. The resulting data are used to set system‐level design inputs for an automated contaminant removal system to process post‐consumer clothing into appropriate feedstocks for recycling. To set the system's levels for automated real‐time near‐infrared analysis, we identify the minimum percentage of primary material that any single garment in a load of used clothing must contain for the average of the full output stream to meet the target purity levels of recyclers. Here, the envisioned automated system can also address undesirable trace materials that might contaminate the processed stream by using imaging cameras coupled with artificial intelligence to identify sections of clothing for de‐trimming. Proof‐of‐concept machine learning algorithms are evaluated to locate and identify trims or garment areas with hidden contaminant materials. Integrating these methods into automated textile cutting systems can provide a cost‐effective means for increasing feedstock purity from used clothing, which can advance circularity for textiles by helping recyclers to reach production volumes and quality targets that were not possible solely with manual dismantling operations.

Parsons, Ryan [Rochester Institute of Technology, ↗

1.2 Mfps standalone X-ray detector for Time-Resolved Experiments

We present a standalone and autonomous X-ray detector capable of operation with the speed of up to1.2Mfps. The detector utilizes UFXC32k hybrid pixel detectors for sensing X-rays, Spartan-6 LX45 FPGA placed in commercially available sbRIO 9628 controller for data acquisition and processing including a compression with zero-suppression algorithm. A Linux-RT system working on the 400 MHz Dual-Core CPU is used for FPGA control and data streaming to the higher-level system over 1 Gbps Ethernet connection. 1.2 M frames per second is achieved in so-called burst mode of operation while in zerodead-time mode 70 kfps is possible. Due to efficient data compression in FPGA there’s no need of using high-speed transceivers and Frame-Grabber cards on the data server side and the detector can stream the data infinitely over standard 1 Gbps network connection. Operation modes were tested at Advanced Photon Source Synchrotron at Argonne National Laboratory.

47 OTHER INSTRUMENTATION↗

Accelerating Advanced Light Source Science Through Multi-Facility HPC Workflows

Synchrotron light sources support a wide array of techniques to investigate materials, often producing complex, high-volume data that challenge traditional workflows. At the Advanced Light Source (ALS), we developed infrastructure to move microtomography data over ESnet to ALCF and NERSC, where CPU- and GPU-based algorithms generate 3D reconstructed volumes of experimental samples. We employ two data movement and reconstruction models: real-time processing as data streams directly to NERSC compute nodes, and automated file transfer to NERSC and ALCF file systems. The streaming pipeline provides users with feedback in under ten seconds, while the file-based workflow produces high-quality reconstructions suitable for deeper analysis in 20-30 minutes. This infrastructure enables users to utilize HPC resources without direct access to backend systems. We plan to extend this architecture to more endstations, supporting our beamline scientists and users.

Abramov, David↗

Hypersonic flows generated by parabolic and paraboloidal shock waves

A computer algorithm has been developed to determine the blunt-body flowfields supporting symmetric parabolic and paraboloidal shock waves at infinite free-stream Mach number. Solutions are expressed in an analytic form as high-order power series, in the coordinate normal to the shock, whose coefficients can be determined exactly. Analytic continuation is provided by the use of Pade approximations. Test cases provide solutions of very high accuracy. In the axisymmetric case for gamma equals 715 the solution has been found far downstream, where it agrees with the modified blast-wave results. For plane flow, on the other hand, a limit line appears within the shock layer, a short distance past the sonic line, suggesting the presence of an imbedded shock. Local solutions in the downstream limit are discussed.

Schwartz, L. W.↗

Advances in Application of Fast Semidirect Computational Methods in Transonic Flow

This paper is intended as a review and summary of the advances made in a recently developed approach for rapid numerical solution of the equations of inviscid transonic aerodynamics. The investigation has been limited to two-dimensional, steady, inviscid flow over airfoils in a subsonic free stream, with emphasis on development of a rapid computational technique, rather than on generality of application. The approach uses finite-difference algorithms called "fast direct elliptic solvers" within an iteration scheme. "Direct" means that the entire computation field is solved at once, rather than in successive traverses over the field as in a point- or line-relaxation method. Such an iterative method is referred to as "semidirect." The iterative convergence can be faster than in other relaxation methods because changes are felt simultaneously at all points in each succeeding iteration. Direct elliptic solvers and semidirect methods have restrictions, but these are gradually being removed. Direct solvers were first developed for solving Poisson's equation on a rectangle without interior boundaries. A method to treat first-order systems, a direct Cauchy-Riemann solver has also been developed. Numerical treatment of part of a system of nonlinear equations by a Poisson solver has been reported. Also Poisson solvers in semidirect methods were used for nonseparable elliptic equations. The semidirect method was extended to the solution of a problem of mixed type, where the improved Murman-Cole transonic small-disturbance difference equations were solved. A slightly supercritical flow over a biconvex airfoil was treated successfully, but the iterations did not converge for more strongly supercritical conditions In another work the addition of terms ot both sides of the difference equations stabilized the iteration for supercritical conditions with large supersonic zones. For this, the Cauchy-Riemann solver was revised to incl,ude the needed terms. Most recently, the evaluation of parameters for rapid convergence and comparisons, with Murman's line-relaxation method was described. The method was extended to full second order accuracy in a fully conservative formulation in another work.

Martin, E. Dale↗

Parallelizing the Unpacking and Clustering of Detector Data for Reconstruction of Charged Particle Tracks on Multi-core CPUs and Many-core GPUs

We present results from parallelizing the unpacking and clustering steps of the raw data from the silicon strip modules for reconstruction of charged particle tracks. Throughput is further improved by concurrently processing multiple events using nested OpenMP parallelism on CPU or CUDA streams on GPU. The new implementation along with earlier work in developing a parallelized and vectorized implementation of the combinatoric Kalman filter algorithm has enabled efficient global reconstruction of the entire event on modern computer architectures. We demonstrate the performance of the new implementation on Intel Xeon and NVIDIA GPU architectures.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Station Keeping of Small Outboard-Powered Boats

Three station keeping controllers have been developed which work to minimize displacement of a small outboard-powered vessel from a desired location. Each of these three controllers has a common initial layer that uses fixed-gain feedback control to calculate the desired heading of the vessel. A second control layer uses a common fixed-gain feedback controller to calculate the net forward thrust, one of two algorithms for controlling engine angle (Fixed-Gain Proportional-integral-derivative (PID) or PID with Adaptively Augmented Gains), and one of two algorithms for differential throttle control (Fixed-Gain PID and PID with Adaptive Differential Throttle gains), which work together to eliminate heading error. The three selected controllers are evaluated using a numerical simulation of a 33-foot center console vessel with twin outboards that is subject to wave, wind, and current disturbances. Each controller is tested for its ability to maintain position in the presence of three sets of environmental disturbances. These algorithms were tested with current velocity of 1.5 m/s, significant wave height of 0.5 m, and wind speeds of 2, 5, and 10 m/s. These values were chosen to model conditions a small vessel may experience in the Gulf Stream off of Fort Lauderdale. The Fixed-gain PID controller progressively got worse as wind speeds increased, while the controllers using adaptive methodologies showed consistent performance over all weather conditions and decreased heading error by as much as 20%. Thus, enhanced robustness to environmental changes has been gained by using an adaptive algorithm.

Fisher, A. D.↗