Engineering PapersSearch

SEARCH · Engineering Papers

Results for “scalable”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

A Scalable Multi-Modal Framework for High-Fidelity Distributed Human Mobility Simulations

The development of data-driven models for human mobility in urban settings requires access to substantial and diverse real-world data. However, existing historical data often presents challenges such as limited volume, variety, and veracity, as well as missing data and privacy preservation concerns. Also, urban mobility modeling is inherently time-variant, complex, and multi-modal, encompassing everything from individual walking and running to private road travel and large-scale public transportation. These challenges call for innovative solutions to overcome data limitations and compute needs to model mobility behaviors accurately. To address these challenges, we propose a distributed, co-simulation-based architecture DURMOSim that integrates real-world data with scalable, high-fidelity simulations, demonstrating distributed co-simulation feasibility with existing mobility models. DURMOSim underpins a modular integration that would enable using any available mobility simulators for greater extensibility and scalability in performing various urban scenarios. In this paper, we present the design, implementation, and performance evaluation of DURMOSim, highlighting its capability to model population-scale mobility patterns. Our initial results show its ability to dynamically synchronize multiple simulation models at runtime with negligible computational overhead. We believe DURMOSim could be a robust tool for advancing urban mobility research and intelligent transportation systems.

Yoginath, Srikanth [ORNL] (ORCID:0000000184236050)

Scalability Analysis of Quantum Models for Stress and Emotion Detection

Stress and emotion detection from high-dimensional physiological signals is a challenging task, particularly when aiming for accurate classification across diverse behavioral states. Quantum machine learning (QML) is promising for modeling such high-dimensional data, but scalability is limited by qubit resources and the exponential cost of classical statevector simulation. This work studies the scalability of quantum support vector machines (QSVMs) for binary stress detection and three-class emotion recognition (Negative/Neutral/Positive) under varying qubit counts and angle-encoding strategies. We also present a comparison study with one-feature-per-qubit (1:1) and two-features-per-qubit (2:1) mappings. Experiments are executed on HPC infrastructure using NVIDIA CUDA-Q to evaluate performance, variance, and class-dependent separability at higher-qubit setups. Results show that larger Hilbert spaces can improve peak accuracy but may increase instability. At the same time, dense 2:1 encoding yields more consistent stress detection performance. For emotion recognition, scaling improves discrimination for classes like Negative and Positive more than Neutral. We find that effective QML scaling is task-dependent and benefits more from encoding design than simply increasing qubit count.

Onim, Md. Saif Hassan [University of Tennessee, Kn

SPADES (Scalable Parallel Discrete Events Simulation) [SWR-24-99]

SPADES (Solver for PArallel Discrete Event Simulation) is an open-source parallel discrete event simulation (PDES) package built on the AMReX library. Targeted at solving discrete event systems in parallel, this software package aims to be performance portable and scalable on heterogeneous computing architectures, e.g., graphic processing units (GPU). SPADES implements optimistic synchronization with rollback through an implementation of the Time Warp algorithm. An alternative conservative synchronization approach is also implemented using the Lower Bound on Incoming Time Stamp. In our implementation, logical processes are represented as cells in a grid and event messages are represented as particles. SPADES supports various parallel decomposition strategies, including the use of the Message Passing Interface (MPI) and OpenMP threading. All major GPU architectures (e.g., Intel, AMD, NVIDIA) are supported through the use of performance portability functionalities implemented in AMReX. The SPADES software is released in NREL Software Record SWR-24-99 “SPADES (Scalable Parallel Discrete Events Simulation)”.

Henry de Frahan, Marc [National Renewable Energy L

Scalable multilevel Monte Carlo methods exploiting parallel redistribution on coarse levels

Here, we study an element agglomeration coarsening strategy that requires data redistribution at coarse levels when the number of coarse elements becomes smaller than the number of MPI processes used on the finest level. The overall procedure generates coarse elements (general unstructured unions of fine grid elements) within the framework of element-based algebraic multigrid methods (or AMGe) studied previously. The AMGe-generated coarse spaces have the ability to exhibit approximation properties of the same order as the fine-level spaces since by construction they contain the piecewise polynomials of the same order as on the fine level. These approximation properties are key for the successful use of AMGe in multilevel solvers for nonlinear partial differential equations as well as for multilevel Monte Carlo (MLMC) simulations. The ability to coarsen without being constrained by the number of MPI processes, as described in the present paper, allows to improve the scalability of these solvers as well as the overall MLMC method. The paper illustrates this latter fact with detailed scalability study of MLMC simulations applied to model Darcy equations with a stochastic log-normal permeability field.

AMGe

Scalable low-loss cryogenic packaging of quantum memories in CMOS-foundry processed photonic chips

Optically linked solid-state quantum memories such as color centers in diamond are a promising platform for distributed quantum information processing and networking. Photonic integrated circuits (PICs) have emerged as a crucial enabling technology for these systems, integrating quantum memories with efficient electrical and optical interfaces in a compact and scalable platform. Packaging these hybrid chips into deployable modules while maintaining low optical loss and resiliency to temperature cycling is a central challenge to their practical use. We demonstrate a packaging method for PICs using surface grating couplers and angle-polished fiber arrays that is robust to temperature cycling, offers scalable channel count, applies to a wide variety of PIC platforms and wavelengths, and offers pathways to automated high-throughput packaging. Using this method, we show optically and electrically packaged quantum memory modules integrating all required qubit controls on chip, operating at millikelvin temperatures with <3 dB losses achievable from fiber to quantum memory for the TE 0 mode at a wavelength of 737 nm.

Bernson, Robert [Tyndall National Institute, Cork

Progress Report for W911NF-23-1-0323: Development of Scalable Three-Dimensional Ion Traps for Quantum Information Processing

Trapped atomic ions are a promising platform for scalable quantum information processing. They lead in many key metrics such as single and two qubit gate fidelities, as well as quantum volume. We recently realized novel 3D ion traps with a high resolution 3D-printing process. The traps combine the scalability potential of surface chip traps with the superior trap performance of macroscopic 3D ion traps. Like photolithography, 3D-printing is a digitally defined process and hence allows for rapid iteration between design and fabrication. In contrast to photolithography, it also allows for fully defined 3D structures and, thus, is an ideal process to develop large, well-defined trap arrays with complex 3D electrode structures. Our current goal is to implement key procedures and perform critical measurements that justify incorporation of 3D-printed traps into large-scale quantum computing efforts. En route, we recently realized trapping in horizontal traps that are operating at cryogenic temperatures.

74 ATOMIC AND MOLECULAR PHYSICS

SMART SiC Power ICs: Scalable, Manufacturable, and Robust Technology for SiC Power Integrated Circuits (Final Technical Report)

This collaborative project was initiated with the goal of developing Scalable, Manufacturable, and Robust Technology for SiC Power Integrated Circuits (SMART SiC Power ICs). In pursuit of this objective, innovative designs and fabrication processes were implemented, enabling the development of large-scale (>1 cm²) SiC Complementary Metal-Oxide-Semiconductor (CMOS) integrated circuits and high-voltage (400–600 V) lateral power MOSFETs (HV-LDMOS) on 150 mm 4H-SiC substrates. The resulting SMART SiC Power ICs are tailored to support a wide range of applications requiring diverse voltage and power levels, including automotive systems, industrial equipment, electronic data processing, energy harvesting, and power conditioning. To achieve the proposed ‘SMART’ technology for SiC ICs, the team focused on 1) the Development of highly scalable CMOS (with high channel mobilities for n-type and p-type MOSFETs), LDMOS (~600V, 10A rated), and IC technologies, 2) Establishment of a manufacturable process baseline in a production-grade-, 150mm, SiC fabrication facility, and 3) Demonstration of SMART SiC ICs. The project initially comprised of fabricating 5 lots. In lot 1 monolithic integration using a single process was achieved. Here, we were able to successfully accomplish Integrated HV NMOSFET with LV CMOS on N-epi/N+ Substrate. The HV NMOS demonstrated a Breakdown Voltage (BV) more than 600V. Circuit demonstration of CMOS was also another achievement from this lot. In lot 2, priority was in place for isolation and integration. Here we addressed the isolation concerns and integrated the HV NMOS and LV CMOS using the N-epi/P-epi/N+ substrate. Similar to the lot 1, we were able to achieve a BV of 600 V for HV NMOS. Optimized gate oxide process with high channel mobilities, better gate oxide reliability, development of SPICE models, successful ohmic process development, novel wafer area saving design layouts, P+ isolation schemes with channeling implantations and high temperature operational circuits demonstrations are some of the key highlights from lot 1 and lot2. In lot 3, discrete device performances of HV NMOS with a BV ~700V and reliable LV CMOS performances were achieved. Also, novel architectural solutions were successfully implemented to suppress the electric field crowding at the gate oxide for reliable operations. In lot 4, half bridge power driver ICs with a conversion efficiency of (target 90% to 95%) in the 1-5MHz switching frequency range for output power between 25 W to 3 kW have been included in. However, due to the unfortunate events of sudden foundry shutdown (SiCamore Semi) the processing of lot 4 wafers came to a complete stop (January 2024). Arrangements have recently been made to shift the fabrication to another foundry, General Electric Aerospace. The fabrication process now on course (as of December 2024). Characterizations are delayed due to this unfortunate circumstance. The proposed trench architectural-based devices and ICs (lot 5) underwent modifications from the original project proposal. This change was necessitated by limitations in the availability of trench-based processes at commercial production-grade fabrication facilities in the US. Apart from above achievements, a Process Development Kit (PDK) was successfully developed for planar type SiC CMOS/LDMOS.

42 ENGINEERING

High-Throughput Microfluidic Electroporation (HTME): A Scalable, 384-Well Platform for Multiplexed Cell Engineering

Electroporation-mediated gene delivery is a cornerstone of synthetic biology, offering several advantages over other methods: higher efficiencies, broader applicability, and simpler sample preparation. Yet, electroporation protocols are often challenging to integrate into highly multiplexed workflows, owing to limitations in their scalability and tunability. These challenges ultimately increase the time and cost per transformation. As a result, rapidly screening genetic libraries, exploring combinatorial designs, or optimizing electroporation parameters requires extensive iterations, consuming large quantities of expensive custom-made DNA and cell lines or primary cells. To address these limitations, we have developed a High-Throughput Microfluidic Electroporation (HTME) platform that includes a 384-well electroporation plate (E-Plate) and control electronics capable of rapidly electroporating all wells in under a minute with individual control of each well. Fabricated using scalable and cost-effective printed-circuit-board (PCB) technology, the E-Plate significantly reduces consumable costs and reagent consumption by operating on nano to microliter volumes. Furthermore, individually addressable wells facilitate rapid exploration of large sets of experimental conditions to optimize electroporation for different cell types and plasmid concentrations/types. Use of the standard 384-well footprint makes the platform easily integrable into automated workflows, thereby enabling end-to-end automation. We demonstrate transformation of E. coli with pUC19 to validate the HTME's core functionality, achieving at least a single colony forming unit in more than 99% of wells and confirming the platform's ability to rapidly perform hundreds of electroporations with customizable conditions. This work highlights the HTME's potential to significantly accelerate synthetic biology Design-Build-Test-Learn (DBTL) cycles by mitigating the transformation/transfection bottleneck.

Gaillard, William R

A Scalable Gaussian Process Approach to Shear Mapping with MuyGPs

Analysis of cosmic shear is an integral part of understanding structure growth across cosmic time, which in turn provides us with information about the nature of dark energy. Conventional methods generate shear maps from which we can infer the matter distribution in the universe. Current methods (e.g., Kaiser–Squires inversion) for generating these maps, however, are tricky to implement and can introduce bias. Recent alternatives construct a spatial process prior for the lensing potential, which allows for inference of the convergence and shear parameters given lensing shear measurements. Realizing these spatial processes, however, scales cubically in the number of observations—an unacceptable expense as near-term surveys expect billions of correlated measurements. Therefore, we present a linearly scaling shear map construction alternative using a scalable Gaussian process prior called MuyGPs. MuyGPs avoids cubic scaling by conditioning interpolation on only nearest neighbors and fits hyperparameters using batched leave-one-out cross-validation. This work is the first step toward a full, scalable mass mapping method. We work in a simplified regime where we validate our method by interpolating and analyzing maps given noisy point-estimate data from all three shear fields, taken from a suite of N -body ray-tracing simulations. We also show that we can perform these operations at the scale of billions of galaxies on high-performance computing platforms.

79 ASTRONOMY AND ASTROPHYSICS

The development of a scalable parallel 3-D CFD algorithm for turbomachinery

Two algorithms capable of computing a transonic 3-D inviscid flow field about rotating machines are considered for parallel implementation. During the study of these algorithms, a significant new method of measuring the performance of parallel algorithms is developed. The theory that supports this new method creates an empirical definition of scalable parallel algorithms that is used to produce quantifiable evidence that a scalable parallel application was developed. The implementation of the parallel application and an automated domain decomposition tool are also discussed.

Luke, Edward Allen

Simple, Scalable, Script-Based Science Processor (S4P)

The development and deployment of data processing systems to process Earth Observing System (EOS) data has proven to be costly and prone to technical and schedule risk. Integration of science algorithms into a robust operational system has been difficult. The core processing system, based on commercial tools, has demonstrated limitations at the rates needed to produce the several terabytes per day for EOS, primarily due to job management overhead. This has motivated an evolution in the EOS Data Information System toward a more distributed one incorporating Science Investigator-led Processing Systems (SIPS). As part of this evolution, the Goddard Earth Sciences Distributed Active Archive Center (GES DAAC) has developed a simplified processing system to accommodate the increased load expected with the advent of reprocessing and launch of a second satellite. This system, the Simple, Scalable, Script-based Science Processor (S42) may also serve as a resource for future SIPS. The current EOSDIS Core System was designed to be general, resulting in a large, complex mix of commercial and custom software. In contrast, many simpler systems, such as the EROS Data Center AVHRR IKM system, rely on a simple directory structure to drive processing, with directories representing different stages of production. The system passes input data to a directory, and the output data is placed in a "downstream" directory. The GES DAAC's Simple Scalable Script-based Science Processing System is based on the latter concept, but with modifications to allow varied science algorithms and improve portability. It uses a factory assembly-line paradigm: when work orders arrive at a station, an executable is run, and output work orders are sent to downstream stations. The stations are implemented as UNIX directories, while work orders are simple ASCII files. The core S4P infrastructure consists of a Perl program called stationmaster, which detects newly arrived work orders and forks a job to run the appropriate executable (registered in a configuration file for that station). Although S4P is written in Perl, the executables associated with a station can be any program that can be run from the command line, i.e., non-interactively. An S4P instance is typically monitored using a simple Graphical User Interface. However, the reliance of S4P on UNIX files and directories also allows visibility into the state of stations and jobs using standard operating system commands, permitting remote monitor/control over low-bandwidth connections. S4P is being used as the foundation for several small- to medium-size systems for data mining, on-demand subsetting, processing of direct broadcast Moderate Resolution Imaging Spectroradiometer (MODIS) data, and Quick-Response MODIS processing. It has also been used to implement a large-scale system to process MODIS Level 1 and Level 2 Standard Products, which will ultimately process close to 2 TB/day.

Lynnes, Christopher

Toward Automatic Scalability Analysis of Message Passing Programs: A Case Study

Scalability analysis forms an important component of any performance debugging cycle, for massively parallel machines. However, tools that help in performing such analysis for parallel programs are non-existent. The primary reason for lack of such tools is the complexity involved in capturing program dynamics such as communication-computation overlap, communication latencies and memory hierarchy reference patterns. In this paper, we highlight some simple techniques that can be used to study scalability of explicit message-passing parallel programs that consider the above issues. We start from the high level source code and use a methodology for deducing communication characteristics and its impact on the total execution time of the program. The approach is validated with the help of a pipelined method for solving scalar tri-diagonal systems, using both simulations and symbolic cost models on the Intel hypercube.

Sarukkai, Sekhar R.

The P-Mesh: A Commodity-based Scalable Network Architecture for Clusters

We designed a new network architecture, the P-Mesh which combines the scalability and fault resilience of a torus with the performance of a switch. We compare the scalability, performance, and cost of the hub, switch, torus, tree, and P-Mesh architectures. The latter three are capable of scaling to thousands of nodes, however, the torus has severe performance limitations with that many processors. The tree and P-Mesh have similar latency, bandwidth, and bisection bandwidth, but the P-Mesh outperforms the switch architecture (a lower bound for tree performance) on 16-node NAB Parallel Benchmark tests by up to 23%, and costs 40% less. Further, the P-Mesh has better fault resilience characteristics. The P-Mesh architecture trades increased management overhead for lower cost, and is a good bridging technology while the price of tree uplinks is expensive.

Nitzberg, Bill

A Scalable Software Architecture Booting and Configuring Nodes in the Whitney Commodity Computing Testbed

The Whitney project is integrating commodity off-the-shelf PC hardware and software technology to build a parallel supercomputer with hundreds to thousands of nodes. To build such a system, one must have a scalable software model, and the installation and maintenance of the system software must be completely automated. We describe the design of an architecture for booting, installing, and configuring nodes in such a system with particular consideration given to scalability and ease of maintenance. This system has been implemented on a 40-node prototype of Whitney and is to be used on the 500 processor Whitney system to be built in 1998.

Fineberg, Samuel A.

A Scalability Model for ECS's Data Server

This report presents in four chapters a model for the scalability analysis of the Data Server subsystem of the Earth Observing System Data and Information System (EOSDIS) Core System (ECS). The model analyzes if the planned architecture of the Data Server will support an increase in the workload with the possible upgrade and/or addition of processors, storage subsystems, and networks. The approaches in the report include a summary of the architecture of ECS's Data server as well as a high level description of the Ingest and Retrieval operations as they relate to ECS's Data Server. This description forms the basis for the development of the scalability model of the data server and the methodology used to solve it.

Menasce, Daniel A.

Validation of a Scalable Solar Sailcraft

The NASA In-Space Propulsion (ISP) program sponsored intensive solar sail technology and systems design, development, and hardware demonstration activities over the past 3 years. Efforts to validate a scalable solar sail system by functional demonstration in relevant environments, together with test-analysis correlation activities on a scalable solar sail system have recently been successfully completed. A review of the program, with descriptions of the design, results of testing, and analytical model validations of component and assembly functional, strength, stiffness, shape, and dynamic behavior are discussed. The scaled performance of the validated system is projected to demonstrate the applicability to flight demonstration and important NASA road-map missions.

Murphy, D. M.

Highly Scalable Matching Pursuit Signal Decomposition Algorithm

Matching Pursuit Decomposition (MPD) is a powerful iterative algorithm for signal decomposition and feature extraction. MPD decomposes any signal into linear combinations of its dictionary elements or atoms . A best fit atom from an arbitrarily defined dictionary is determined through cross-correlation. The selected atom is subtracted from the signal and this procedure is repeated on the residual in the subsequent iterations until a stopping criterion is met. The reconstructed signal reveals the waveform structure of the original signal. However, a sufficiently large dictionary is required for an accurate reconstruction; this in return increases the computational burden of the algorithm, thus limiting its applicability and level of adoption. The purpose of this research is to improve the scalability and performance of the classical MPD algorithm. Correlation thresholds were defined to prune insignificant atoms from the dictionary. The Coarse-Fine Grids and Multiple Atom Extraction techniques were proposed to decrease the computational burden of the algorithm. The Coarse-Fine Grids method enabled the approximation and refinement of the parameters for the best fit atom. The ability to extract multiple atoms within a single iteration enhanced the effectiveness and efficiency of each iteration. These improvements were implemented to produce an improved Matching Pursuit Decomposition algorithm entitled MPD++. Disparate signal decomposition applications may require a particular emphasis of accuracy or computational efficiency. The prominence of the key signal features required for the proper signal classification dictates the level of accuracy necessary in the decomposition. The MPD++ algorithm may be easily adapted to accommodate the imposed requirements. Certain feature extraction applications may require rapid signal decomposition. The full potential of MPD++ may be utilized to produce incredible performance gains while extracting only slightly less energy than the standard algorithm. When the utmost accuracy must be achieved, the modified algorithm extracts atoms more conservatively but still exhibits computational gains over classical MPD. The MPD++ algorithm was demonstrated using an over-complete dictionary on real life data. Computational times were reduced by factors of 1.9 and 44 for the emphases of accuracy and performance, respectively. The modified algorithm extracted similar amounts of energy compared to classical MPD. The degree of the improvement in computational time depends on the complexity of the data, the initialization parameters, and the breadth of the dictionary. The results of the research confirm that the three modifications successfully improved the scalability and computational efficiency of the MPD algorithm. Correlation Thresholding decreased the time complexity by reducing the dictionary size. Multiple Atom Extraction also reduced the time complexity by decreasing the number of iterations required for a stopping criterion to be reached. The Course-Fine Grids technique enabled complicated atoms with numerous variable parameters to be effectively represented in the dictionary. Due to the nature of the three proposed modifications, they are capable of being stacked and have cumulative effects on the reduction of the time complexity.

Christensen, Daniel

Three-dimensional Finite Element Formulation and Scalable Domain Decomposition for High Fidelity Rotor Dynamic Analysis

This paper has two objectives. The first objective is to formulate a 3-dimensional Finite Element Model for the dynamic analysis of helicopter rotor blades. The second objective is to implement and analyze a dual-primal iterative substructuring based Krylov solver, that is parallel and scalable, for the solution of the 3-D FEM analysis. The numerical and parallel scalability of the solver is studied using two prototype problems - one for ideal hover (symmetric) and one for a transient forward flight (non-symmetric) - both carried out on up to 48 processors. In both hover and forward flight conditions, a perfect linear speed-up is observed, for a given problem size, up to the point of substructure optimality. Substructure optimality and the linear parallel speed-up range are both shown to depend on the problem size as well as on the selection of the coarse problem. With a larger problem size, linear speed-up is restored up to the new substructure optimality. The solver also scales with problem size - even though this conclusion is premature given the small prototype grids considered in this study.

Datta, Anubhav