Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “layout”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Rapid Prototyping Techniques for Organic Direct-Bonded-Copper Power Modules

Organic Direct-Bonded-Copper (ODBC) is a novel packaging technology for power modules which allows higher flexibility in layout design. In this work, a set of rapid prototyping techniques are developed for ODBC modules based on a polyimide dielectric material. These techniques enable fast and low-cost fabrication of modules with 3-Dimensional (3D) layout features. An example half-bridge (HB) silicon carbide (SiC) metal oxide semiconductor field effect transistor (MOSFET) module is designed and prototyped using the proposed techniques aiming at ultra-low power loop inductance. Finite element analysis (FEA)-circuit co-simulations results and experimental results validate an approximately 0.71nH power loop inductance for the example module design.

36 MATERIALS SCIENCE↗

Performance Analysis of Data Processing in Distributed File Systems with Near Data Processing

In the era of big data, the escalating volume and velocity of data generation pose significant challenges in data processing. Traditional systems like Spark and Hadoop manage the increasing amount and velocity of data by improving data placement and processing speeds. However, they face inherent limitations due to the essential data movement required for processing. In this paper, we explore the Skyhook framework, a novel extension of the Ceph distributed system, which significantly reduces the need for data movement. We present an extensive case study using the Skyhook framework, applying it with the TPC-H and K-means clustering algorithms. More specifically, we leverage the TPC-H benchmark to distinguish between CPU-intensive and I/O-intensive tasks. We explore the integration of K-means clustering into SQL, coupled with a near-data processing system to offload the computational burden of the K-means clustering algorithm to storage nodes. We conduct a comprehensive performance evaluation of distributed data processing applications across three processing approaches: traditional layout (baseline), optimized layout, and near-data processing. Additionally, we introduce the use of the FIO tool to simulate real-world system workloads, enabling the measurement of performance metrics such as average latency and CPU utilization. Our research is a significant advance in understanding how to optimize data processing systems to meet the demands of the modern data landscape.

Hou, Shiyue↗

Methodology of Low Inductance Busbar Design for Three-Level Converters

Three-level (3L) converters are more susceptible to parasitics compared with two-level converters because of their complicated structure with multiple switching loops. In this article, the methodology of busbar layout design for 3L converters based on the magnetic cancellation effect is presented. The methodology can fit for 3L converters with symmetric and asymmetric configurations. A detailed design example is provided for a high-power 3L-active neutral point clamped (ANPC) converter, which includes the module selection, busbar layout, and dc-link capacitor placement. The loop inductance of the busbar is verified with simulation, impedance measurements, and converter experiments. Furthermore, the results match with each other, and the inductances of short and long loops are 6.5 and 17.5 nH, respectively, which are significantly lower than the busbars of NPC-type converters in other references.

42 ENGINEERING↗

Concept Lens: Visual Comparison and Evaluation of Generative Model Manipulations

Generative models are becoming a transformative technology for the creation and editing of images. However, it remains challenging to harness these models for precise image manipulation. These challenges often manifest as inconsistency in the editing process, where both the type and amount of semantic change, depend on the image being manipulated. Moreover, there exist many methods for computing image manipulations, whose development is hindered by the matter of inconsistency. This paper aims to address these challenges by improving how we evaluate, compare, and explore the space of manipulations offered by a generative model. We present Concept Lens, a visual interface that is designed to aid users in understanding semantic concepts carried in image manipulations, and how these manipulations vary over generated images. Given the large space of possible images produced by a generative model, Concept Lens is designed to support the exploration of both generated images, and their manipulations, at multiple levels of detail. To this end, the layout of Concept Lens is informed by two hierarchies: a hierarchical organization of (1) original images, grouped by their similarities, and (2) image manipulations, where manipulations that induce similar changes are grouped together. This layout allows one to discover the types of images that consistently respond to a group of manipulations, and vice versa, manipulations that consistently respond to a group of codes. We show the benefits of this design across multiple use cases, specifically, studying the quality of manipulations for a single method, and offering a means of comparing different methods.

clustering↗

Reimagining Disassembly Interfaces With Visualization: Combining Instruction Tracing and Control Flow With DisViz

In applications where efficiency is critical, developers may examine their compiled binaries, seeking to understand how the compiler transformed their source code and what performance implications that transformation may have. This analysis is challenging due to the vast number of disassembled binary instructions and the many-to-many mappings between them and the source code. These problems are exacerbated as source code size increases, giving the compiler more freedom to map and disperse binary instructions across the disassembly space. Interfaces for disassembly typically display instructions as an unstructured listing or sacrifice the order of execution. Here, we design a new visual interface for disassembly code that combines execution order with control flow structure, enabling analysts to both trace through code and identify familiar aspects of the computation. Central to our approach is a novel layout of instructions grouped into basic blocks that displays a looping structure in an intuitive way. We add to this disassembly representation a unique block-based mini-map that leverages our layout and shows context across thousands of disassembly instructions. Finally, we embed our disassembly visualization in a web-based tool, DisViz, which adds dynamic linking with source code across the entire application. DizViz was developed in collaboration with program analysis experts following design study methodology and was validated through evaluation sessions with ten participants from four institutions. Participants successfully completed the evaluation tasks, hypothesized about compiler optimizations, and noted the utility of our new disassembly view. Our evaluation suggests that our new integrated view helps application developers in understanding and navigating disassembly code.

Computer science↗

Analysis and Optimization of a Multi-Layer Integrated Organic Substrate for High Current GaN HEMT-Based Power Module

In this paper, analysis and optimization of a multi-layer organic substrate for high current GaN HEMT based power module are discussed. The organic multi-layer substrates can provide high electrical performance in terms of low parasitic inductance in the power loop by providing vertical layout, and shielding for reduction of common-mode noise, a common problem in fast switching power converters. Furthermore, high performance cooling solutions, such as micro-channel heat sinks, can be directly bonded to the substrate for optimum thermal management. The structure of the proposed architecture, thermal analysis and optimization of layer thickness, thermo-mechanical stress analysis of the GaN HEMT and development of a high-performance heat sink are discussed.

47 OTHER INSTRUMENTATION↗

Status Quo of Heliostat Field Deployment Processes

Deployment of the solar field of a concentrating solar power plant is one of many factors that are integral to the success of a project. Knowledge transfer from outside the industry is limited due to the unique nature of heliostats, which redirect sunlight to a receiver with high precision while maintaining a high level of reflectivity. Moreover, learning from project to project can be limited due to the site-specific nature of projects, as the market includes several developers, each with their own unique design. In this paper, we discuss the state of the art in heliostat field deployment. We cover all the key aspects of deployment from project assessment to a fully functioning system, which include site selection, layout development, supply chain, assembly, site preparation and construction, calibration, and operations and maintenance.

concentrating solar power↗

Feasibility of crystal-assisted collimation in the CERN accelerator complex

Bent silicon crystals mounted on high-accuracy angular actuators were installed in the CERN Super Proton Synchrotron (SPS) and extensively tested to assess the feasibility of crystal-assisted collimation in circular hadron colliders. The adopted layout was exploited and regularly upgraded for about a decade by the UA9 Collaboration. The investigations provided the compelling evidence of a strong reduction of beam losses induced by nuclear inelastic interactions in the aligned crystals in comparison with amorphous orientation. A conceptually similar device, installed in the betatron cleaning insertion of CERN Large Hadron Collider (LHC), was operated through the complete acceleration and storage cycle and demonstrated a large reduction of the background leaking from the collimation region and radiated into the cold sections of the accelerator and the experimental detectors. The implemented layout and the relevant results of the beam tests performed in the SPS and in the LHC with stored proton and ion beams are extensively discussed.

43 PARTICLE ACCELERATORS↗

Accelerating Random Forest Classification on GPU and FPGA

Random Forests (RFs) are a commonly used machine learning method for classification and regression tasks spanning a variety of application domains, including bioinformatics, business analytics, and software optimization. While prior work has focused primarily on improving performance of the training of RFs, many applications, such as malware identification, cancer prediction, and banking fraud detection, require fast RF classification. In this work, we accelerate RF classification on GPU and FPGA. In order to provide efficient support for large datasets, we propose a hierarchical memory layout suitable to the GPU/FPGA memory hierarchy. We design three RF classification code variants based on that layout, and we investigate GPU- and FPGA-specific considerations for these kernels. Our experimental evaluation, performed on an Nvidia Xp GPU and on a Xilinx Alveo U250 FPGA accelerator card using publicly available datasets on the scale of millions of samples and tens of features, covers various aspects. First, we evaluate the performance benefits of our hierarchical data structure over the standard compressed sparse row (CSR) format. Second, we compare our GPU implementation with cuML, a machine learning library targeting Nvidia GPUs. Third, we explore the performance/accuracy tradeoff resulting from the use of different tree depths in the RF. Finally, we perform a comparative performance analysis of our GPU and FPGA implementations. Our evaluation shows that for high accuracy targets, our GPU implementation yields 5-9x speedup over CSR, and up to a 2x speedup over cuML.

FPGA, Xilinx FPGA, GPU, Random Forest classificati↗

Realistic Cost to Execute Practical Quantum Circuits using Direct Clifford+T Lattice Surgery Compilation

We report a resource estimation pipeline that explicitly compiles quantum circuits expressed using the Clifford+T gate set into a surface code lattice surgery instruction set. The cadence of magic state requests from the compiled circuit enables the optimization of magic state distillation and storage requirements in a post-hoc analysis. To compile logical circuits into lattice surgery operations, we build upon the open-source Lattice Surgery Compiler. The revised compiler operates in two stages: the first translates logical gates into an abstract, layout-independent instruction set; the second compiles these into local lattice surgery instructions that are allocated to hardware tiles according to a specified resource layout. The second stage retains logical parallelism while avoiding resource contention in the fault-tolerant layer, aiding realism. Additionally, users can specify dedicated tiles at which magic states are replenished, enabling resource costs from the logical computation to be considered independently from magic state distillation and storage. We demonstrate the applicability of our pipeline to large practical quantum circuits by providing resource estimates for the ground state estimation of molecules. Finally, we find that variable magic state consumption rates in real circuits can cause the resource costs of magic state storage to dominate unless production is varied to suit.

97 MATHEMATICS AND COMPUTING↗

MemFriend: Understanding Memory Performance with Spatial-Temporal Affinity

In HPC applications, memory access behavior is one of the main factors affecting performance. Improving an application’s memory access behavior involves optimizing data layout and/or restructuring code, and requires studying spatial-temporal data locality. Existing data locality analyses focus on single-location metrics and are restricted to evaluating temporal locality. We introduce spatial-temporal affinity metrics that quantify temporal access proximity, forward access correlation, and nearby access correlation between pairs of memory locations. We describe methods for distinguishing between potential vs. realized affinity and for reasoning about affinity at multiple resolutions (3D, 2D, 1D). Finally, we construct spatial-temporal affinity signatures that classify memory behavior and that be used to reason about changes in software (data relayout, code refactoring) or hardware (caching, prefetching). We describe methods for signature visualization, interpretation, and quantitative comparison of signatures. We evaluate our methodology using applications with variants that contrast data structures, data layouts and algorithms. We show that spatial-temporal affinity analysis provides novel insights and enables predictive reasoning about application performance when contrasted with reuse distance analysis.

Suriyakumar, Yasodhadevi↗

Exploration ToolKit (ExTK)

The Exploration Toolkit (ExTK) is a reusable Extended Reality (XR) system developed for incorporating and exploring 3-Dimensional (3-D) computer aided design (CAD) models in XR, with a primary focus on Augmented Reality (AR). The ExTK consists of a Developer Mode and a User Mode. In Developer Mode, ExTK provides developers with the ability to easily import 3-D CAD models and activate desired exploration functionality and layout. Multiple models can be added to a single instantiation of the ExTK using Unity's Scene capability. Exploration functions include scaling, rotating, explode/contract, animations, hiding parts, submodules, and measurement functions. In User Mode, ExTK provides a menu system that allows users to select models and initiate exploration functions. ExTK is architected for reusability and developers can customize the ExTK layout and functions according to application needs. ExTK is designed to be hardware agnostic, although initial development focused on the Microsoft HoloLens as the primary deployment platform. The ExTK is developed using the Unity Game Engine Development Platform. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525. SAND2021-1506 O

Klein, BrandonThorin↗

Experimental study of airpath electrification in an opposed-piston two stroke (OP2S) engine architecture

The opposed-piston two stroke (OP2S) engine shows potential as an alternative engine architecture to the conventional four stroke engine due to its high-power density, thermal efficiency, and versatile airpath management system. Since the pistons of a two-stroke engine do not pump the air into and out of the cylinder like in a four-stroke engine, the selection of the air induction devices and airpath actuators becomes critical to optimize engine performance. Both the pumping losses and the in-cylinder combustion process can be affected by the scavenging process in a two-stroke engine. Therefore, this study compares two different airpath configurations for the same family of OP2S engines and investigates performance metrics like scavenging control, pumping work, net indicated and brake efficiencies, and engine-out emissions associated with each airpath. Data was collected on a 3.2 L, two-cylinder OP2S engine with an electrically assisted turbocharger (EAT) and a 4.9 L displacement, three-cylinder engine with a variable geometry turbocharger (VGT) and a supercharger. The experiments consisted of speed and load sweeps for both engines at the same operating conditions to compare scavenge control in both architectures. For the three-cylinder layout, the SE sweep range was much higher, and the intake pressure could be independently varied with air flowrate, thus providing more flexibility for scavenging control. The supercharger and the VGT usage was optimized based on its efficiency map and thus, this layout had lower pumping losses compared to the EAT. The two-cylinder engine had a higher overall SE as compared to the three-cylinder engine, but the intake pressure and air flowrate could not be decoupled, leading to over scavenging and increased short circuiting of fresh charge into the exhaust.

Bhatt, Ankur [Clemson University, Clemson, SC, USA↗

Earthquake response of head-mounted equipment in advanced nuclear reactors

The seismic response of safety-related equipment mounted on the head of an advanced reactor, including pumps, control rod drive mechanisms, and reactor monitoring devices, will affect the design and layout of many advanced reactors. High earthquake-induced accelerations in such equipment may challenge their seismic qualification and trigger the need for additional support framing on the reactor head. Base isolation is a design solution that can drastically reduce seismic demands on equipment. This article describes a set of earthquake-simulator experiments conducted on a scale-model of a base-isolated reactor vessel including four representations of head-mounted equipment, with frequencies spanning from 4.5 to 27 Hz. Dynamic responses of the head-mounted equipment, including displacements, accelerations, and strains, were measured in the experiments for three support conditions: conventional, and seismically isolated using single concave Friction Pendulum (SFP) bearings and triple Friction Pendulum (TFP) bearings. Seismic isolation was effective at reducing equipment responses (accelerations, displacements, and strains) with respect to those in the conventionally supported vessel across a range of seismic inputs. Companion numerical studies highlight the accuracy to be expected in the calculation of different response quantities for lightly damped equipment. The importance of characterizing damping in head-mounted, safety-related equipment through physical experiments to support design and risk assessment is made clear through the numerical simulations.

Engineering↗

Two-stage formation-energy correction (NbZr, TaZr, VZr)

This bundle contains the scripts, the raw and corrected per-structure data, and the manuscript plots for the NbZr / TaZr / VZr BCC binary formation energies and the associated RMSDs. Why a two-stage correction is necessary: The "raw" formation energy of every relaxed VASP configuration is computed in the usual way, FE_raw(c) = E_alloy(c) - sum_i x_i * E_pure_i , where E_pure_i are the per-atom total energies of the pure-element reference structures (Nb, Ta, V, Zr in the same BCC supercell, with identical INCAR / KPOINTS / PAW choices). With perfectly consistent reference runs the raw FE should vanish at the two pure-element endpoints (x = 0 and x = 1) by construction. In practice this does not hold for two reasons that are present in our dataset: 1. Reference-energy inconsistency (composition-dependent bias). Even with identical input parameters, the pure-element runs (stored in `corrected_DFT_pure_element_runs/`) differ slightly from the values that would be implied by the alloy runs at near-pure compositions (a few meV/atom). This bias is approximately linear in concentration, because the residual error in E_pure_Nb (or E_pure_Ta / E_pure_V) propagates into FE_raw(c) as (1 - x) * dE_pure_1, and the corresponding error in E_pure_Zr propagates as x * dE_pure_2. Left uncorrected, this produces a non-physical "tilt" of FE_raw(x) and shifts the entire FE-vs-x cloud away from zero at the endpoints. 2. Endpoint anchoring against the audited true endpoints. The strict endpoint values (FE_x0_meVatom, FE_x1_meVatom in `corrected_fe_strict_endpoints_20260518/strict_endpoint_check_20260518.csv`) were re-derived from an independent cross-check of the pure-element runs. After stage 1 removes the linear bias, the near-pure compositions in the alloy dataset still extrapolate to values that differ slightly from these audited endpoints — because stage 1 is fit from a few near-end alloy bins, not from the audited pure-element references themselves. The README.txt file discusses how these issues are addressed by the two-stage correction, and describes folder layout, pipeline summary, and how to re-run.

36 MATERIALS SCIENCE↗

CONCEPT OF A POLARIZED POSITRON SOURCE FOR CEBAF

Positron beams would provide new and meaningful probes for the experimental program at the Thomas Jefferson National Accelerator Facility (JLab), including but not limited to future hadronic physics and dark matter experiments. Critical requirements involve generating positron beams with a high degree of spin polarization, sufficient intensity and a continuous-wave (CW) bunch train compatible with acceleration to 12 GeV at the Continuous Electron Beam Accelerator Facility (CEBAF). To address these requirements, a polarized positron injector based upon the bremsstrahlung of an intense CW spin polarized electron beam is considered*. First a polarized electron beam line provides >1 mA of polarized electrons at ~120 MeV to a high-power target for positron production. Next, a second beam line collects, shapes and aligns the spin of positrons for users. Finally, the positron beam is matched into the CEBAF acceptance for acceleration and transport to the end stations with energies up to 12 GeV. An optimized layout to provide positrons beams with intensity >100 nA (polarized) or intensity >3 µA (unpolarized) will be discussed in this poster.

Habet, S. H.↗

Positron beams at Ce+BAF

Positron beams would provide a new and meaningful probe for the experimental program at the Thomas Jefferson National Accelerator Facility (JLab). The JLab Positron Working Group, formed in 2018 and now with over 250 members from 75 institutions, continues to develop an experimental program with high duty-cycle positron beams including but not limited to future hadronic physics and dark matter experiments. Critical requirements involve generating positron beams with a high degree of spin polarization, sufficient intensity and a continuous-wave (CW) bunch train compatible with acceleration to 12 GeV at the Continuous Electron Beam Accelerator Facility (CEBAF). In this presentation we describe a start-to-end layout for positron beams at 12 GeV CEBAF utilizing the Low Energy Research Facility (LERF) at Jefferson Lab to build two new injectors. A GaAs dc high voltage photo-gun first generates >1 mA of polarized electrons which are then accelerated to 80-150 MeV and directed to a high-power spinning W target for polarized bremsstrahlung and positron pair creation. A second injector then collects, bunches and accelerates the positrons to 123 MeV. The positron beams are transported by a new beam line and injected into the CEBAF acceptance for acceleration to the end stations with energies up to 12 GeV. The layout is optimized to provide Users with positron spin polarization >60% and intensity greater than >100 nA, and with higher intensities when polarization is not required.

Benesch, J.↗

bifacial_radiance: a python package for modeling bifacial solar photovoltaic systems

bifacial_radiance is a national-laboratory-developed, community-supported, open-source toolkit that provides a set of functions and classes for simulating the performance of bifacial photovoltaic (PV) systems. (Bifacial PV modules collect light on the front as well as the rear side.) bifacial_radiance automates calculations of PV system layout and performance to use along with the popular ray-tracing software tool RADIANCE (Ward, 1994). Specific algorithms include design and layout of PV modules, reflective ground surfaces, shading obstructions, and irradiance calculations throughout the system, among others. bifacial_radiance is an important component of a growing ecosystem of open-source tools for solar energy (William F Holmgren et al., 2018).

97 MATHEMATICS AND COMPUTING↗