Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “streaming algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

A System to Provide Deterministic Flight Software Operation and Maximize Multicore Processing Performance: The Safe and Precise Landing – Integrated Capabilities Evolution (SPLICE) Datapath

A method and design are described for a system that processes multiple data streams, utilizing a multicore asymmetric processing architecture, that eliminates data interrupts to the application processors. The design supports a deterministic environment for flight software in NASA’s Safe and Precise Landing – Integrated Capabilities Evolution (SPLICE) project. The SPLICE project develops sensor, algorithm, and compute technologies for Precision Landing and Hazard Avoidance (PL&HA) capabilities. The compute technology for SPLICE is the Descent and Landing Computer (DLC). The DLC hosts several SPLICE algorithms with high computational resource requirements that must be executed in a real-time and deterministic manner. The software runs on a custom Single Board Computer (SBC), with a Xilinx Ultrascale+ Multiprocessor System-on-a-Chip (MPSoC). Input data for the flight software is from a variety of sensors, unique with respect to data rate and packet size. A data path between the SPLICE sensors and algorithms is designed to efficiently deliver this data to the flight software using the MPSoC asymmetric processing cores and Field Programmable Gate Array (FPGA) fabric. This is implemented in a manner that isolates the application processors running the flight software from interrupts associated with the input data. By leveraging real-time processors on the MPSoC, and a structure with the appropriate interfaces in the shared memory on the SBC, the flight software can use the full set of application processors. The available utilization for each processor in this set is also maximized for the SPLICE applications, providing a sufficiently deterministic execution environment without the cost and overhead of a real-time operating system.

David K. Rutishauser↗

Development of an explicit multigrid algorithm for quasi-three-dimensional viscous flows in turbo-machinery

A rapid quasi three-dimensional analysis was developed for blade-to-blade flows in turbomachinery. The analysis solves the unsteady Euler or thin layer Navier-Stokes equations in a body-fitted coordinate system. It accounts for the effects of rotation, radius change, and stream-surface thickness. The Baldwin-Lomax eddy-viscosity model is used for turbulent flows. The equations which are solved by a two-stage Runge-Kutta scheme made efficient by use of vectorization, a variable time-step, and a flux-based multigrid scheme, are described. A stability analysis is presented for the two-stage scheme. Results for a flat-plate model problem show the applicability of the method to axial, radial, and rotating geometries. Results for a centrifugal impeller and a radial diffuser show that the quasi three-dimensional viscous analysis can be a practical design tool.

Chima, R. V.↗

Real-time processing of radar return on a parallel computer

NASA is working with the FAA to demonstrate the feasibility of pulse Doppler radar as a candidate airborne sensor to detect low altitude windshears. The need to provide the pilot with timely information about possible hazards has motivated a demand for real-time processing of a radar return. Investigated here is parallel processing as a means of accommodating the high data rates required. A PC based parallel computer, called the transputer, is used to investigate issues in real time concurrent processing of radar signals. A transputer network is made up of an array of single instruction stream processors that can be networked in a variety of ways. They are easily reconfigured and software development is largely independent of the particular network topology. The performance of the transputer is evaluated in light of the computational requirements. A number of algorithms have been implemented on the transputers in OCCAM, a language specially designed for parallel processing. These include signal processing algorithms such as the Fast Fourier Transform (FFT), pulse-pair, and autoregressive modelling, as well as routing software to support concurrency. The most computationally intensive task is estimating the spectrum. Two approaches have been taken on this problem, the first and most conventional of which is to use the FFT. By using table look-ups for the basis function and other optimizing techniques, an algorithm has been developed that is sufficient for real time. The other approach is to model the signal as an autoregressive process and estimate the spectrum based on the model coefficients. This technique is attractive because it does not suffer from the spectral leakage problem inherent in the FFT. Benchmark tests indicate that autoregressive modeling is feasible in real time.

Aalfs, David D.↗

Optimum Transonic Airfoils Based on the Euler Equations

We solve the problem of determining airfoils that approximate, in a least square sense, given surface pressure distributions in transonic flight regimes. The flow is modeled by means of the Euler equations and the solution procedure is an adjoint- based minimization algorithm that makes use of the inverse Theodorsen transform in order to parameterize the airfoil. Fast convergence to the optimal solution is obtained by means of the pseudo-time method. Results are obtained using three different pressure distributions for several free stream conditions. The airfoils obtained have given a trailing edge angle.

Iollo, Angelo↗

Sage: parallel semi-asymmtric graph algorithms for NVRAMs

Non-volatile main memory (NVRAM) technologies provide an attractive set of features for large-scale graph analytics, including byte-addressability, low idle power, and improved memory-density. NVRAM systems today have an order of magnitude more NVRAM than traditional memory (DRAM). NVRAM systems could therefore potentially allow very large graph problems to be solved on a single machine, at a modest cost. However, a significant challenge in achieving high performance is in accounting for the fact that NVRAM writes can be much more expensive than NVRAM reads. In this paper, we propose an approach to parallel graph analytics using the Parallel Semi-Asymmetric Model (PSAM), in which the graph is stored as a read-only data structure (in NVRAM), and the amount of mutable memory is kept proportional to the number of vertices. Similar to the popular semi-external and semi-streaming models for graph analytics, the PSAM approach assumes that the vertices of the graph fit in a fast read-write memory (DRAM), but the edges do not. In NVRAM systems, our approach eliminates writes to the NVRAM, among other benefits. To experimentally study this new setting, we develop Sage, a parallel semi-asymmetric graph engine with which we implement provably-efficient (and often work-optimal) PSAM algorithms for over a dozen fundamental graph problems. We experimentally study Sage using a 48--core machine on the largest publicly-available real-world graph (the Hyperlink Web graph with over 3.5 billion vertices and 128 billion edges) equipped with Optane DC Persistent Memory, and show that Sage outperforms the fastest prior systems designed for NVRAM. Importantly, we also show that Sage nearly matches the fastest prior systems running solely in DRAM, by effectively hiding the costs of repeatedly accessing NVRAM versus DRAM.

97 MATHEMATICS AND COMPUTING↗

Understanding the Impact of Data Staging for Coupled Scientific Workflows

We report the rate of data generated by cutting-edge experimental science facilities and large-scale simulations enabled by current high-performance computing (HPC) systems has continued to grow at a far greater pace than the development of the network and storage capabilities on which these systems rely. To cope with this challenge, scientist are moving toward the creation of autonomous experiments and HPC simulations using machine learning. However, efficiently moving, storing, and processing large amounts of data away from the point of origin presents an incredible challenge. In-memory computing, in situ analysis, data staging, and data streaming are recognized viable alternatives to traditional file-based methods for transferring data between coupled workflows. However, the performance trade-offs and limitations for these methods are not fully understood when used in HPC applications. This article presents a comprehensive performance assessment of the current solutions for data staging when applied to applications that are not necessary I/O intensive which makes them not ideal candidates for these methods. Our study is based on experiments running at scale on Oak Ridge National Laboratory's Summit supercomputer using applications and simulations that cover typical computational motifs and patterns. We investigated the usability and cost/benefit trade-offs of staging algorithms for HPC applications under different scenarios and highlight opportunities for optimizing the dataflow between coupled simulation workflows.

97 MATHEMATICS AND COMPUTING↗

Data processing in infrared astronomy

Infrared astronomy is often carried out with rocket probes or orbiting satellite telescopes in order to escape the effects of atmospheric absorption. The data returned from such missions is a highly abstracted digital representation of measurements made by analog detectors. The ability to extract infrared-emission information from these data streams depends on a thorough understanding of the information flow from the telescope aperture to the computer center. This paper reviews the primary elements of this end-to-end concept and the impact of each of these elements on the data processing algorithms, including the division between onboard and ground processing for scientific measurements.

Pelzmann, R. F., Jr.↗

Signal Processing Based Method for Real-Time Anomaly Detection in High-Performance Computing

Performance anomalies can manifest as irregular execution times or abnormal execution events for many reasons, including network congestion and resource contention. Detecting such anomalies in real-time by analyzing the details of performance traces at scale is impractical due to the sheer volume of data High-Performance Computing (HPC) applications produce. In this paper, we propose formulating HPC performance anomaly detection as a signal-processing problem where anomalies can be treated as noise. We evaluate our proposed method in comparison with two other commonly used anomaly detection techniques of varying complexity based on their detection accuracy and scalability. Since real-time in-situ anomaly detection at a large scale requires lightweight methods that can handle a large volume of streaming data, we find that our proposed method provides the best trade-off. We then implement the proposed method in Chimbuko, the first online, distributed, and scalable workflow-level performance trace analysis framework. We compare our proposed signal-based anomaly detection algorithm with two other methods using a function of their accuracy, F1 score, and detection overhead. Our experiments demonstrate that our proposed approach achieves a 99% improvement for the benchmark datasets and a 93% improvement with Chimbuko traces.

99 GENERAL AND MISCELLANEOUS↗

Resource Allocation for Single Carrier Massive MIMO Systems

Resource allocation in orthogonal frequency division multiplexing (OFDM) systems is performed through allocating blocks of subcarriers to each user. Even though OFDM is the primary waveform for 5G NR systems, research reports have noted that single carrier modulation (SCM) offers several advantages over OFDM in massive multiple input multiple output (MIMO) systems, making it a preferred candidate for some future applications such as massive machine type communications (mMTC). This paper presents a method for SCM resource allocation and the relevant information recovery algorithms at the receiver. Our emphasis is on cyclic prefixed SCM, where highly flexible and efficient frequency domain detection algorithms enable the operation of many simultaneous users in a massive MIMO uplink scenario. The proposed resource allocation method allows the number of users to exceed the number of antennas at the base station (BS). Each single carrier transmission is partitioned into L interleaved streams, and each user is allocated a number of such streams. One major benefit of SCM is that each data symbol is spread over the entire bandwidth. As such, the receiver performance is dictated by the average channel gain across the transmission band rather than the channel gain at a given frequency bin or a small group of frequencies. In the proposed setup, each stream may be thought of as a resource block in SCM, analogous to resource blocks in OFDM. Hence, in the context of this paper, the terms resource blocks and streams may be used interchangeably.

5G and Beyond Communications↗

Edge detection algorithm for SST images

An algorithm to detect fronts in satellite-derived sea surface temperature fields is presented. Although edge detection is the main focus, the problem of cloud detection is also addressed since unidentified clouds can lead to erroneous edge detection. The algorithm relies on a combination of methods and it operates at the picture, the window, and the local level. The resulting edge detection is not based on the absolute strength of the front, but on the relative strength depending on the context, thus, making the edge detection temperature-scale invariant. The performance of this algorithm is shown to be superior to that of simpler algorithms commonly used to locate edges in satellite-derived SST images. This evaluation was performed through a careful comparison between the location of the fronts obtained by applying the various methods to the SST images and the in situ measures of the Gulf Stream position.

Cayula, Jean-Francois↗

Compensating For Unbalance In Pulse-Code Phase Modulation

Algorithm proposed for use in pulse-code phase-modulation transmitter in which non-return-to-zero (NRZ) or biphase data modulated directly onto radio-frequency residual carrier signal. Devised to compensate somewhat for effect, upon distant receiver, of unbalance in stream of transmitted data. Formulated to compute combinations of modulation index, data rate, and transmitter power compensating for measured unbalance in transmitted data stream.

Nguyen, Tien M.↗

Artificial Boundary Conditions for Computation of Oscillating External Flows

In this paper, we propose a new technique for the numerical treatment of external flow problems with oscillatory behavior of the solution in time. Specifically, we consider the case of unbounded compressible viscous plane flow past a finite body (airfoil). Oscillations of the flow in time may be caused by the time-periodic injection of fluid into the boundary layer, which in accordance with experimental data, may essentially increase the performance of the airfoil. To conduct the actual computations, we have to somehow restrict the original unbounded domain, that is, to introduce an artificial (external) boundary and to further consider only a finite computational domain. Consequently, we will need to formulate some artificial boundary conditions (ABC's) at the introduced external boundary. The ABC's we are aiming to obtain must meet a fundamental requirement. One should be able to uniquely complement the solution calculated inside the finite computational domain to its infinite exterior so that the original problem is solved within the desired accuracy. Our construction of such ABC's for oscillating flows is based on an essential assumption: the Navier-Stokes equations can be linearized in the far field against the free-stream back- ground. To actually compute the ABC's, we represent the far-field solution as a Fourier series in time and then apply the Difference Potentials Method (DPM) of V. S. Ryaben'kii. This paper contains a general theoretical description of the algorithm for setting the DPM-based ABC's for time-periodic external flows. Based on our experience in implementing analogous ABC's for steady-state problems (a simpler case), we expect that these boundary conditions will become an effective tool for constructing robust numerical methods to calculate oscillatory flows.

Tsynkov, S. V.↗

Sampling Technique for Robust Odorant Detection Based on MIT RealNose Data

This technique enhances the detection capability of the autonomous Real-Nose system from MIT to detect odorants and their concentrations in noisy and transient environments. The lowcost, portable system with low power consumption will operate at high speed and is suited for unmanned and remotely operated long-life applications. A deterministic mathematical model was developed to detect odorants and calculate their concentration in noisy environments. Real data from MIT's NanoNose was examined, from which a signal conditioning technique was proposed to enable robust odorant detection for the RealNose system. Its sensitivity can reach to sub-part-per-billion (sub-ppb). A Space Invariant Independent Component Analysis (SPICA) algorithm was developed to deal with non-linear mixing that is an over-complete case, and it is used as a preprocessing step to recover the original odorant sources for detection. This approach, combined with the Cascade Error Projection (CEP) Neural Network algorithm, was used to perform odorant identification. Signal conditioning is used to identify potential processing windows to enable robust detection for autonomous systems. So far, the software has been developed and evaluated with current data sets provided by the MIT team. However, continuous data streams are made available where even the occurrence of a new odorant is unannounced and needs to be noticed by the system autonomously before its unambiguous detection. The challenge for the software is to be able to separate the potential valid signal from the odorant and from the noisy transition region when the odorant is just introduced.

Duong, Tuan A.↗

Development of a Global Reference Surface Reflectance and BRDF Datasets from Geostationary Satellite Observations and AERONET Measurements

Surface reflectances and their dependency on illumination-view geometries (i.e., BRDF) are the foundation of many high-level satellite products for land and water monitoring. Yet it is difficult to evaluate the quality of satellite-based surface reflectances with ground-based measurements due to the spatial scale differences. In order to fill the gap, here we develop a reference dataset of surface reflectance and BRDF at the global AERONET sites with data streams from operational geostationary sensors including Himawari 8/9 AHI, GK-2A AMI, and GOES 16/17 ABI. Taking the top-of-atmosphere (TOA) reflectance and the site measured atmospheric aerosol optical depth (AOD) as the main inputs, we apply the GeoNEX-AC algorithm to performance accurate atmospheric correction and derive 10-minute surface reflectance and daily Ross-Thick-Li-Sparse (RTLS) BRDF parameters at AERONET sites where coincident AOD measurements and TOA observations are available from 2016 (for Himawari) or 2018 (for GOES) onwards. The algorithm ensures that the retrieved surface BRDF parameters, along with the site-measured AOD, allow the atmospheric radiative transfer model, SHARM, accurately simulate the observed TOA reflectance at diurnal and longer time scales. They are our best estimates of the surface optical properties and thus can serve as the “reference” to evaluate the performance of operational atmospheric correction algorithms (where AOD is assumed unknown and needs to be retrieved). The reference BRDF also allow us to evaluate the spectral band ratios between the SWIR (e.g., 2200 nm) and the visible (e.g., 650 nm) regions, which are commonly used in operational atmospheric correction algorithms. Finally, we demonstrate that the reference dataset can be used to develop potential data synergies between different GEO satellites as well as GEO-LEO sensors.

Weile Wang↗

Robust clustering of the local Milky Way stellar kinematic substructures with Gaia eDR3

Understanding local stellar kinematic substructures in the solar neighbourhood helps build a complete picture of the formation of the Milky Way, as well as an empirical phase space distribution of dark matter that would inform detection experiments. We apply the clustering algorithm HDBSCAN on the Gaia early third data release to identify a list of stable clusters in velocity space and action-angle space by taking into account the measurement uncertainties and studying the stability of the clustering results. We find 1405 (497) stars in 23 (6) robust clusters in velocity space (action-angle space) that are consistently not associated with noise. We discuss the kinematic properties of these structures and study whether many of the small clusters belong to a similar larger cluster based on their chemical abundances. They are attributed to the known structures: the Gaia Sausage-Enceladus, the Helmi Stream, and globular cluster NGC 3201 are found in both spaces, while NGC 104 and the thick disc (Sequoia) are identified in velocity space (action-angle space). Although we do not identify any new structures, we find that the HDBSCAN member selection of already known structures is unstable to input kinematics of the stars when resampled within their uncertainties. We therefore present the stable subset of local kinematic structures, which are consistently identified by the clustering algorithm, and emphasize the need to take into account error propagation during both the manual and automated identification of stellar structures, both for existing ones as well as future discoveries.

79 ASTRONOMY AND ASTROPHYSICS↗

Relaxation and Preconditioning for High Order Discontinuous Galerkin Methods with Applications to Aeroacoustics and High Speed Flows

This project is about the investigation of the development of the discontinuous Galerkin finite element methods, for general geometry and triangulations, for solving convection dominated problems, with applications to aeroacoustics. Other related issues in high order WENO finite difference and finite volume methods have also been investigated. methods are two classes of high order, high resolution methods suitable for convection dominated simulations with possible discontinuous or sharp gradient solutions. In [18], we first review these two classes of methods, pointing out their similarities and differences in algorithm formulation, theoretical properties, implementation issues, applicability, and relative advantages. We then present some quantitative comparisons of the third order finite volume WENO methods and discontinuous Galerkin methods for a series of test problems to assess their relative merits in accuracy and CPU timing. In [3], we review the development of the Runge-Kutta discontinuous Galerkin (RKDG) methods for non-linear convection-dominated problems. These robust and accurate methods have made their way into the main stream of computational fluid dynamics and are quickly finding use in a wide variety of applications. They combine a special class of Runge-Kutta time discretizations, that allows the method to be non-linearly stable regardless of its accuracy, with a finite element space discretization by discontinuous approximations, that incorporates the ideas of numerical fluxes and slope limiters coined during the remarkable development of the high-resolution finite difference and finite volume schemes. The resulting RKDG methods are stable, high-order accurate, and highly parallelizable schemes that can easily handle complicated geometries and boundary conditions. We review the theoretical and algorithmic aspects of these methods and show several applications including nonlinear conservation laws, the compressible and incompressible Navier-Stokes equations, and Hamilton-Jacobi-like equations.

Shu, Chi-Wang↗

Spontaneous Raman–LIF–CO–OH measurements of species concentration in turbulent spray flames

This paper presents new measurements of species concentrations, temperature and mixture fraction in selected regions of a turbulent ethanol spray flame. The line-Raman–LIF–COsingle bondOH setup developed at the Sandia's Combustion Research Facility is utilised to probe regions of a spray flame where laser breakdown of liquid droplets is avoided and the remaining interferences can be corrected. The spray flame is stabilised on the piloted Sydney needle spray burner, where axial translation of the liquid injecting needle in the air-blast stream can transition the spray from dilute to dense. The solution to obtaining successful measurements is found to be multifaceted and includes: the appropriate selection of flame conditions; high sensitivity of the Raman detection system permitting reduced laser energies; development of a pre-processing algorithm to reject strong droplet interferences; and application of the hybrid matrix inversion method combined with wavelet denoising to account for interference corrections and noise at the very low signal levels obtained. Unique and necessary for the successful measurements reported in this paper, a pre-processing algorithm is outlined that removes data points corrupted with strong interferences from droplets. These interferences arise from a range of sources, but the most intense are due to the laser interaction with surrounding mist or liquid fragments, such that measurements near the jet centreline are corrupted and hence discarded. Reliable measurements of mixture fraction, temperature obtained from the sum of the species number densities, and species mole fractions are reported for regions in the flames sufficiently far from the centreline. The paper demonstrates the feasibility of the judicious use of Raman scattering in turbulent spray flames, the results of which will be extremely useful for validating numerical simulations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Nearby stellar substructures in the Galactic halo from DESI Milky Way Survey Year 1 Data Release

We report five nearby ($d_{\mathrm{helio}} < 5$ kpc) stellar substructures in the Galactic halo from a subset of 138 661 stars in the Dark Energy Spectroscopic Instrument (DESI) Milky Way Survey Year 1 Data Release. With an unsupervised clustering algorithm, HDBSCAN*, these substructures are independently identified in Integrals of Motion ($E_{\rm tot}$, $L_{\rm z}$, $\log {J_r}$, $\log {J_z}$) space and Galactocentric cylindrical velocity space ($V_{R}$, $V_{\phi }$, $V_{z}$). We associate all identified clusters with known nearby substructures (Helmi streams, M18-Cand10/MMH-1, Sequoia, Antaeus, and ED-2) previously reported in various studies. With metallicities precisely measured by DESI, we confirm that the Helmi streams, M18-Cand10, and ED-2 are chemically distinct from local halo stars. We have characterized the chemodynamic properties of each dynamic group, including their metallicity dispersions, to associate them with their progenitor types (globular cluster or dwarf galaxy). Our approach for searching substructures with HDBSCAN* reliably detects real substructures in the Galactic halo, suggesting that applying the same method can lead to the discovery of new substructures in future DESI data. With more stars from future DESI data releases and improved astrometry from the upcoming Gaia Data Release 4, we will have a more detailed blueprint of the Galactic halo, offering a significant improvement in our understanding of the formation and evolutionary history of the Milky Way Galaxy.

dynamics↗