Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “UPS”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Efficient Data Management in Neutron Scattering Data Reduction Workflows at ORNL

Oak Ridge National Laboratory (ORNL) experimental neutron science facilities produce 1.2 TB a day of raw event-based data that is stored using the standard metadata-rich NeXus schema built on top of the HDF5 file format. Performance of several data reduction workflows is largely determined by the amount of time spent on the loading and processing algorithms in Mantid, an open-source data analysis framework used across several neutron sciences facilities around the world. The present work introduces new data management algorithms to address identified input output (I/O) bottlenecks on Mantid. First, we introduce an in-memory binary-tree metadata index that resemble NeXus data access patterns to provide a scalable search and extraction mechanism. Second, data encapsulation in Mantid algorithms is optimally redesigned to reduce the total compute and memory runtime footprint associated with metadata I/O reconstruction tasks. Results from this work show speed ups in wall-clock time on ORNL data reduction workflows, ranging from 11% to 30% depending on the complexity of the targeted instrument-specific data. Nevertheless, we highlight the need for more research to address reduction challenges as experimental data volumes increase.

Godoy, William↗

Efficient loading of reduced data ensembles produced at ORNL SNS/HFIR neutron time-of-flight facilities

We present algorithmic improvements to the loading operations of certain reduced data ensembles produced from neutron scattering experiments at Oak Ridge National Laboratory (ORNL) facilities. Ensembles from multiple measurements are required to cover a wide range of the phase space of a sample material of interest. They are stored using the standard NeXus schema on individual HDF5 files. This makes it a scalability challenge, as the number of experiments stored increases in a single ensemble file. The present work follows up on our previous efforts on data management algorithms, to address identified input output (I/O) bottlenecks in Mantid, an open-source data analysis framework used across several neutron science facilities around the world. We reuse an in-memory binary-tree metadata index that resembles data access patterns, to provide a scalable search and extraction mechanism. In addition, several memory operations are refactored and optimized for the current common use cases, ranging most frequently from 10 to 180, and up to 360 separate measurement configurations. Results from this work show consistent speed ups in wall-clock time on the Mantid LoadMD routine, ranging from 19% to 23% on average, on ORNL production computing systems. The latter depends on the complexity of the targeted instrument-specific data and the system I/O and compute variability for the shared computational resources available to users of ORNL’s Spallation Neutron Source (SNS) and the High Flux Isotope Reactor (HFIR) instruments. Nevertheless, we continue to highlight the need for more research to address reduction challenges as experimental data volumes, user time and processing costs increase.

Godoy, William↗

LOGAN: High-Performance GPU-Based X-Drop Long-Read Alignment

Pairwise sequence alignment is one of the most computationally intensive kernels in genomic data analysis, accounting for more than 90% of the runtime for key bioinformatics applications. This method is particularly expensive for third-generation sequences due to the high computational cost of analyzing sequences of length between 1Kb and 1Mb. Given the quadratic overhead of exact pairwise algorithms for long alignments, the community primarily relies on approximate algorithms that search only for high-quality alignments and stop early when one is not found. In this work, we present the first GPU optimization of the popular X-drop alignment algorithm, that we named LOGAN. Results show that our high-performance multi-GPU implementation achieves up to 181.6 GCUPS and speed-ups up to 6.6× and 30.7× using 1 and 6 NVIDIA Tesla V100, respectively, over the state-of-the-art software running on two IBM Power9 processors using 168 CPU threads, with equivalent accuracy. We also demonstrate a 2.3× LOGAN speed-up versus ksw2, a state-of-art vectorized algorithm for sequence alignment implemented in minimap2, a long-read mapping software. Furthermore, to highlight the impact of our work on a real-world application, we couple LOGAN with a many-to-many long-read alignment software called BELLA, and demonstrate that our implementation improves the overall BELLA runtime by up to 10.6×. Finally, we adapt the Roofline model for LOGAN and demonstrate that our implementation is near optimal on the NVIDIA Tesla V100s.

97 MATHEMATICS AND COMPUTING↗

FPGA-Accelerated Range-Limited Molecular Dynamics

Long timescale Molecular Dynamics (MD) simulation of small molecules is crucial in drug design and basic science. To accelerate a small data set that is executed for a large number of iterations, high-efficiency is required. Recent work in this domain has demonstrated that among COTS devices only FPGA-centric clusters can scale beyond a few processors. The problem addressed here is that, as the number of on-chip processors has increased from fewer than 10 into the hundreds, previous intra-chip routing solutions are no longer viable. We find, however, that through various design innovations, high efficiency can be maintained. These include replacing the previous broadcast networks with ring-routing and then augmenting the rings with out-of-order and caching mechanisms. Others are adding a level of hierarchical filtering and memory recycling. Two novel optimized architectures emerge, together with a number of variations. These are validated, analyzed, and evaluated. We find that in the domain of interest speed-ups over GPUs are achieved. Finally, the potential impact is that this system promises to be the basis for scalable long timescale MD with commodity clusters.

97 MATHEMATICS AND COMPUTING↗

Machine Learning-Based Extreme Data Reduction for Prompt Supernova Pointing at DUNE

One of the goals of the Deep Underground Neutrino Experiment (DUNE) is to use the massive underground liquid argon time projection chamber (LArTPC) detectors at its far site for multimessenger astronomy (MMA), in the detection of neutrinos from core-collapse supernovae (SNe). Its current baseline trigger strategy detects activity in the detector that is consistent with supernova (SN) neutrinos and saves the raw data for further offline analysis but provides no prompt pointing information crucial for optical follow-ups by other observatories. This approach is based on the assumption that prompt pointing determination using raw data is computationally prohibitive. In this article, we demonstrate a proof-of-concept based on applying extreme data reduction on the buffered SN data in the DUNE data acquisition (DAQ) system’s front-end computers using a machine learning (ML) workflow. This reduces the data by ~5 orders of magnitude, allowing a full track reconstruction to be carried out quickly on a single server. The total time to perform the ML-based data reduction and the full track reconstruction is less than the time to transfer the SN data back to Fermilab or a high-performance computing (HPC) center. This shows that prompt processing of raw SN data is possible and, in fact, trivial once the data have been reduced to reject radiological backgrounds, paving the way to a high-quality SN pointing trigger that is based on fully reconstructed data instead of trigger primitives (TPs).

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

MAGIC: M arching Cubes Isosurface Uncertainty Visualization for G auss i an Uncertain Data With Spatial C orrelation

Here, in this paper, we study the propagation of data uncertainty through the marching cubes algorithm for isosurface visualization for correlated uncertain data. Consideration of correlation has been shown paramount for avoiding errors in uncertainty quantification and visualization in multiple prior studies. Although the problem of isosurface uncertainty with spatial data correlation has been previously addressed, there are two major limitations to prior treatments. First, there are no analytical formulations for uncertainty quantification of isosurfaces when the data uncertainty is characterized by a Gaussian distribution with spatial correlation. Second, as a consequence of the lack of analytical formulations,existing techniques resort to a Monte Carlo sampling approach, which is expensive and difficult to integrate into visualization tools. To address these limitations, we present a closed-form framework to efficiently derive uncertainty in marching cubes level-sets for Gaussian uncertain data with spatial correlation (MAGIC). To derive closed-form solutions, we leverage the Hinkley's derivation on the ratio of Gaussian distributions. With our analytical framework, we achieve a significant speed-up and enhanced accuracy of uncertainty quantification over classical Monte Carlo methods. We further accelerate our analytical solutions using many-core processors to achieve speed-ups up to 585× and integrability with production visualization tools for broader impact. We demonstrate the effectiveness of our correlation-aware uncertainty framework through experiments on meteorology, urban flow, and astrophysics simulation datasets.

Gaussian↗

Identification of Dynamic Force Coefficients for an Additively Manufactured Hermetic Squeeze Film Bearing Support Damper Utilizing a Pass-Through Channel

Abstract The following paper presents breakthrough experimental results for a new hermetic squeeze film damper (HSFD) concept that is integrally designed within an externally pressurized tilting-pad radial gas bearing support. The flexibly damped gas bearing module was designed for a 7.2″ (183 mm) diameter shaft and fabricated using direct metal laser melting (DMLM); also known as additive manufacturing. The bearing and HSFD were sized based on ongoing studies for oil-free supercritical carbon dioxide (sCO2) power turbines in the 8.5 MW–10 MW power range. The development of the new damper concept was motivated by past dynamic testing on HSFD, which generated frequency-dependent stiffness and damping force coefficients. In efforts to eliminate the frequency dependency, a new HSFD architecture was conceived that adds accumulator volumes and a pass-through channel to previously conceived HSFD flow network designs. The other motivation for the work is the need to develop a cost-effective and reliable oil-free bearing technology that is scalable to large power turbomachinery applications. There were several objectives for the following work. The first objective was to successfully design and fabricate a single piece bearing-damper using additive manufacturing, while dimensionally controlling critical design features. The paper discusses the manufacturing steps and shows cut-ups that reveal adequate clearance control capability with internal damper clearances. The second objective was to perform experimental testing with the new HSFD design in efforts to extract stiffness and damping coefficients for excitation frequencies within 20–160 Hz and peak vibration amplitudes between 0.25 mils (6.35 microns) to 1 mil (25.4 microns). The test results for a single HSFD bearing module indicated that the design modifications to the HSFD architecture were successful in eliminating nearly all the frequency dependencies for the stiffness force coefficient. The dynamic tests yielded a stiffness coefficient that varied between 112 klb/in. (19.6 MN/m) and 96 klb/in. (16.8 MN/m). The damping force coefficient however, exhibited relatively more variation with frequency with values residing between 175 lb-s/in. (31 kN-s/m) to 214 lb-s/in. (37 kN-s/m). Finally, the paper advances a three-dimensional fluid-structure interaction (FSI) model using transient finite element analysis (FEA) coupled to a computational fluid dynamics (CFD) model. The FSI analysis performed between 20 Hz and 80 Hz was used to predict the stiffness and damping of the HSFD using a quarter-section model of the damper. The FSI analysis was able to support test results by showing only a 6–7.4% change in the magnitude of force coefficients. Stiffness predictions agree reasonably well with experiments whereas damping is underpredicted.

Engineering↗

Fast Multigrid Reduction-in-Time for Advection via Modified Semi-Lagrangian Coarse-Grid Operators

Many iterative parallel-in-time algorithms have been shown to be highly efficient for diffusion-dominated partial differential equations (PDEs) but are inefficient or even divergent when applied to advection-dominated PDEs. We consider the application of the multigrid reduction-in-time (MGRIT) algorithm to linear advection PDEs. Here, the key to efficient time integration with this method is using a coarse-grid operator that provides a sufficiently accurate approximation to the so-called ideal coarse-grid operator. For certain classes of semi-Lagrangian discretizations, we present a novel semi-Lagrangian-based coarse-grid operator that leads to fast and scalable multilevel time integration of linear advection PDEs. The coarse-grid operator is composed of a semi-Lagrangian discretization followed by a correction term, with the correction designed so that the leading-order truncation error of the composite operator is approximately equal to that of the ideal coarse-grid operator. Parallel results show substantial speed-ups over sequential time integration for variable-wave-speed advection problems in one and two spatial dimensions, and using high-order discretizations up to order five. The proposed approach establishes the first practical method that provides small and scalable MGRIT iteration counts for advection problems.

97 MATHEMATICS AND COMPUTING↗

Randomized Algorithms for Symmetric Nonnegative Matrix Factorization

Symmetric Nonnegative Matrix Factorization (SymNMF) is a technique in data analysis and machine learning that approximates a matrix with a product of a nonnegative, low-rank matrix and it transpose. To design faster and more scalable algorithms for SymNMF we develop two randomized algorithms for its computation. The first method uses randomized matrix sketching to compute an initial low-rank approximation to the input matrix and proceeds to uses this as a low-rank input to rapidly compute a SymNMF. The second methods uses randomized leverage score sampling to approximately solve constrained least squares problems. Many successful methods for SymNMF rely on (approximately) solving sequences of constrained least squares problems. Here, we prove theoretically that leverage score sampling can approximately solve constrained least squares problems to e-accuracy. Finally we demonstrate both methods work in practice by applying them to graph clustering tasks on large real world data sets. These experiments show that our methods approximately maintain solution quality and achieve significant speed ups for both large dense and large sparse problems.

97 MATHEMATICS AND COMPUTING↗

Analysis Background & Noise in Stretched Wire Alignment Technique Measurements

The Stretched-Wire Alignment Technique (SWAT) is one method of magnet alignment for linear induction accelerators. The applications of SWAT have been implemented for aligning solenoid magnets on the Scorpius linear induction accelerator which will be sited at the Nevada National Security Site and the Flash X-Ray (FXR) linear induction accelerator at Lawrence Livermore National Laboratory’s Contained Firing Facility. This article describes both systematic (repeatable) and random sources of background and noise as well as practical ways to eliminate or reduce them to acceptable levels. Systematic sources include reflections from wire ends, rapid sag due to ohmic heating of the wire, magnetic materials, and shot rate. Random sources include air currents, vibration of nearby equipment, mechanical stability of test equipment, and the instruments used to measure the wire motion. Mitigations include curve fitting and adaptive noise signal cancellation, and mechanical damping. Finite Element Analysis (FEA) was used to identify and resolve a repeatable wire vibration frequency interfering with the signal resolution. Two stretched wire alignment technique set ups from Sandia National Labs and Lawrence Livermore National Lab have shown background noise sources and ways of mitigating them by either analysis methods or change of mechanical configuration. Conclusions that were drawn included the severe sensitivity of the deflection to even small external interferences of the SWAT wire such that it requires attention to detail in mechanical set up and analysis.

Linear Inductive Accelerator↗

Curb Allocation and Pick-Up Drop-Off Aggregation for a Shared Autonomous Vehicle Fleet

Advances in information technologies and vehicle automation have birthed new transportation services, including shared autonomous vehicles (SAVs). Shared autonomous vehicles are on-demand self-driving taxis, with flexible routes and schedules, able to replace personal vehicles for many trips in the near future. The siting and density of pick-up and drop-off (PUDO) points for SAVs, much like bus stops, can be key in planning SAV fleet operations, since PUDOs impact SAV demand, route choices, passenger wait times, and network congestion. Unlike traditional human-driven taxis and ride-hailing vehicles like Lyft and Uber, SAVs are unlikely to engage in quasi-legal procedures, like double parking or fire hydrant pick-ups. In congested settings, like central business districts (CBD) or airport curbs, SAVs and others will not be allowed to pick up and drop off passengers wherever they like. This paper uses an agent-based simulation to model the impact of different PUDO locations and densities in the Austin, Texas CBD, where land values are highest and curb spaces are coveted. In this paper 18 scenarios were tested, varying PUDO density, fleet size and fare price. The results show that for a given fare price and fleet size, PUDO spacing (e.g., one block vs. three blocks) has significant impact on ridership, vehicle-miles travelled, vehicle occupancy, and revenue. A good fleet size to serve the region’s 80 core square miles is 4000 SAVs, charging a $1 fare per mile of travel distance, and with PUDOs spaced three blocks of distance apart from each other in the CBD.

Hunter, Christian B.↗

Multiscale X-ray phase contrast imaging of human cartilage for investigating osteoarthritis formation

The evolution of cartilage degeneration is still not fully understood, partly due to its thinness, low radio-opacity and therefore lack of adequately resolving imaging techniques. X-ray phase-contrast imaging (X-PCI) offers increased sensitivity with respect to standard radiography and CT allowing an enhanced visibility of adjoining, low density structures with an almost histological image resolution. This study examined the feasibility of X-PCI for high-resolution (sub-) micrometer analysis of different stages in tissue degeneration of human cartilage samples and compare it to histology and transmission electron microscopy. Ten 10%-formalin preserved healthy and moderately degenerated osteochondral samples, post-mortem extracted from human knee joints, were examined using four different X-PCI tomographic set-ups using synchrotron radiation the European Synchrotron Radiation Facility (France) and the Swiss Light Source (Switzerland). Volumetric datasets were acquired with voxel sizes between 0.7 × 0.7 × 0.7 and 0.1 × 0.1 × 0.1 µm 3 . Data were reconstructed by a filtered back-projection algorithm, post-processed by ImageJ, the WEKA machine learning pixel classification tool and VGStudio max. For correlation, osteochondral samples were processed for histology and transmission electron microscopy. X-PCI provides a three-dimensional visualization of healthy and moderately degenerated cartilage samples down to a (sub-)cellular level with good correlation to histologic and transmission electron microscopy images. X-PCI is able to resolve the three layers and the architectural organization of cartilage including changes in chondrocyte cell morphology, chondrocyte subgroup distribution and (re-)organization as well as its subtle matrix structures. X-PCI captures comprehensive cartilage tissue transformation in its environment and might serve as a tissue-preserving, staining-free and volumetric virtual histology tool for examining and chronicling cartilage behavior in basic research/laboratory experiments of cartilage disease evolution.

3D analysis↗

Performance improvements of the windowed multipole formalism using a rational fraction approximation of the Faddeeva function

The windowed multipole (WMP) formalism was introduced as a way to calculate Doppler broadened cross sections on the fly during Monte Carlo simulations. While more arithmetic is needed compared to point-wise cross section look-ups, performance remained competitive from the large memory reductions and sequential data access. The single most expensive function call in a depleted fuel assembly problem using WMP comes from the evaluation of the Faddeeva function, which previously relied on a highly accurate, highly-branching algorithm. This paper explores the use of rational fraction approximations tailored to the domain interest of reactor physics applications and the development of lower accuracy approximations sufficient for our application. The rational approximations were implemented and tested in OpenMC on an infinite medium problem to stress the cross section calculation routine and a PWR assembly problem. In both cases, the rational approximation nearly eliminated the ∼ 20% penalty previously observed when comparing to point-wise libraries. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Analysis of the performance of a hybrid CPU/GPU 1D2D coupled model for real flood cases

Coupled 1D2D models emerged as an efficient solution for a two-dimensional (2D) representation of the floodplain combined with a fast one-dimensional (1D) schematization of the main channel. At the same time, high-performance computing (HPC) has appeared as an efficient tool for model acceleration. In this work, a previously validated 1D2D Central Processing Unit (CPU) model is combined with an HPC technique for fast and accurate flood simulation. Due to the speed of 1D schemes, a hybrid CPU/GPU model that runs the 1D main channel on CPU and accelerates the 2D floodplain with a Graphics Processing Unit (GPU) is presented. Since the data transfer between sub-domains and devices (CPU/GPU) may be the main potential drawback of this architecture, the test cases are selected to carry out a careful time analysis. Here, the results reveal the speed-up dependency on the 2D mesh, the event to be solved and the 1D discretization of the main channel. Additionally, special attention must be paid to the time step size computation shared between sub-models. In spite of the use of a hybrid CPU/GPU implementation, high speed-ups are accomplished in some cases.

54 ENVIRONMENTAL SCIENCES↗

How To Manual - 4.56

The "how to" document is designed to help walk the analyst through difficult aspects of software usage. It should supplement both the User's manual and the Theory document, by providing examples and detailed discussion that reduce learning time for complex set ups. These documents are intended to be used together. We will not formally list all parameters for an input here — see the User's manual for this. All the examples in the "How To" document are part of the Sierra/SD test suite, and each will run with no modification. The nature of this document casts together a number of rather unrelated procedures. Grouping them is difficult. Please try to use the table of contents and the index as a guide in finding the analyses of interest.

97 MATHEMATICS AND COMPUTING↗

Sierra/SD-- How To Manual - 4.58

The “how to” document is designed to help walk the analyst through difficult aspects of software usage. It should supplement both the User’s manual and the Theory document, by providing examples and detailed discussion that reduce learning time for complex set ups. These documents are intended to be used together. We will not formally list all parameters for an input here – see the User’s manual for this. All the examples in the “How To” document are part of the Sierra/SD test suite, and each will run with no modification. The nature of this document casts together a number of rather unrelated procedures. Grouping them is difficult. Please try to use the table of contents and the index as a guide in finding the analyses of interest.

97 MATHEMATICS AND COMPUTING↗

Sierra/SD - How To Manual, 5.0

The “how to” document guides the user through complicated aspects of software usage. It should supplement both the User’s manual and the Theory document, by providing examples and detailed discussion that reduce learning time for complex set ups. These documents are intended to be used together. We will not formally list all parameters for an input here – see the User’s manual for this. All the examples in the “How To” document are part of the Sierra/SD test suite, and each will run with no modification. The nature of this document casts together a number of rather unrelated procedures. Grouping them is difficult. Please try to use the table of contents and the index as a guide in finding the analyses of interest.

97 MATHEMATICS AND COMPUTING↗