Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Linux”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

MultiPhATE2: code for functional annotation and comparison of phage genomes

To address a need for improved tools for annotation and comparative genomics of bacteriophage genomes, we developed multiPhATE2. As an extension of multiPhATE, a functional annotation code released previously, multiPhATE2 performs gene finding using multiple algorithms, compares the results of the algorithms, performs functional annotation of coding sequences, and incorporates additional search algorithms and databases to extend the search space of the original code. MultiPhATE2 performs gene matching among sets of closely related bacteriophage genomes, and uses multiprocessing to speed computations. MultiPhATE2 can be re-started at multiple points within the workflow to allow the user to examine intermediate results and adjust the subsequent computations accordingly. In addition, multiPhATE2 accommodates custom gene calls and sequence databases, again adding flexibility. MultiPhATE2 was implemented in Python 3.7 and runs as a command-line code under Linux or MAC operating systems. Full documentation is provided as a README file and a Wiki website.

59 BASIC BIOLOGICAL SCIENCES↗

AXEAP : a software package for X-ray emission data analysis using unsupervised machine learning

The Argonne X-ray Emission Analysis Package ( AXEAP ) has been developed to calibrate and process X-ray emission spectroscopy (XES) data collected with a two-dimensional (2D) position-sensitive detector. AXEAP is designed to convert a 2D XES image into an XES spectrum in real time using both calculations and unsupervised machine learning. AXEAP is capable of making this transformation at a rate similar to data collection, allowing real-time comparisons during data collection, reducing the amount of data stored from gigabyte-sized image files to kilobyte-sized text files. With a user-friendly interface, AXEAP includes data processing for non-resonant and resonant XES images from multiple edges and elements. AXEAP is written in MATLAB and can run on common operating systems, including Linux, Windows, and MacOS.

97 MATHEMATICS AND COMPUTING↗

Gaia: segmented germanium detector for high-energy X-ray fluorescence and spectroscopic imaging

We present Gaia, a monolithic array of 96 high-purity germanium pixel detectors integrated with a custom low-noise application-specific integrated circuit (ASIC) and a field-programmable gate array (FPGA)-based data acquisition system. The sensor operates at ∼100 K using a commercial closed-cycle cryocooler, with the in-vacuum electronics thermally isolated from the cold finger to ensure thermal stability. The system demonstrates an average energy resolution of 711 eV at 122 keV, measured using a 57 Co source, and 253 eV at 5.89 keV, measured with 55 Fe across all channels. The readout architecture incorporates a high-performance FPGA paired with a dual-core ARM processor, forming a complete embedded Linux-based computing platform. Communication between the processor and FPGA is handled via memory-mapped I/O, and data are streamed over high-speed gigabit Ethernet. A full-scale 384-pixel Gaia detector, based on this 96-element module, is currently under fabrication.

36 MATERIALS SCIENCE↗

Open-Source Architecture for Multi-Party Update Verification for Data Acquisition Devices

Power grids are integral parts of modern daily life and are increasingly under cyberattack. Common software update processes for grid control devices rely on only one organization to provide verification. This paper roposes a multi-signature software update process to help better secure data acquisition devices from malicious actions by using standard cryptosystems such as TLS. A prototype system build on Linux and off-the-shelf hardware has shown successful update and attack prevention with our customized multi-party update software.

Newberg, Benjamin↗

Quantum/AI Topology-Aware Latency-Adaptive HPC Workflow Scheduling Optimization

The growing demand for more powerful high-performance computing (HPC) systems has led to a steady rise in energy consumption by supercomputing worldwide. This study is focused on comparing our Application-Topology Mapper (ATMapper) to the popular Simple Linux Utility for Resource Management (SLURM) for the purpose of exploring methods that can further optimize job-scheduling within HPC systems. ATMapper is an Artificial-Intelligence based approach to job-scheduling that is currently being enhanced with quantum annealing (QA) to generate optimal schedules faster. We are applying QA to speedup our ATMapper process to achieve higher computing efficiency, thereby reducing HPC energy consumption. Here, we examine how four job-scheduling approaches perform in processor node assignment when using an example network architecture of 4 interconnected nodes. Using a specialized script, we are assessing the schedule of a computation flow with 11 interdependent tasks. The data movements among nodes were tracked to count for the number of interactions (network hops) between nodes needed to complete the tasks. The total number of hops and the job completion time were then used to quantify the efficiency of the different mapping approaches. In addition to SLURM, we also compare our ATMapper to the QA-enabled LBNL TIGER and the D-Wave Distributed Computing processor assignment approaches. The preliminary results showed that our topology-aware, latency-adaptive ATMapper is significantly more efficient when compared to the other scheduling approaches due to its load-imbalance network allocation. The scheduler displayed a computing efficiency of 53% by performing significantly fewer network hops than its alternatives. By reducing the number of hops, ATMapper was able to perform all 11 tasks by using only 3 nodes out of given 4. This research indicates the potential to use QA/AI for HPC job-scheduling. Later, we will test a SLURM simulator program to draw further comparisons on the effectiveness of ATMapper's scheduling approach. The results of this comparison will serve as a baseline for later improving SLURM's performance using a QA-enhanced ATMapper approach.

Caraveo, Braulio [University of Huston - Clear Lak↗

Towards Improving Container Security by Preventing Runtime Escapes

Container escapes enable the adversary to execute code on the host from inside an isolated container. Notably, these high severity escape vulnerabilities originate from three sources: (1) container profile misconfigurations, (2) Linux kernel bugs, and (3) container runtime vulnerabilities. While the first two cases have been studied in the literature, no works have investigated the impact of container runtime vulnerabilities. In this paper, to fill this gap, we study 59 CVEs for 11 different container runtimes. As a result of our study, we found that five of the 11 runtimes had nine publicly available PoC container escape exploits covering 13 CVEs. Our further analysis revealed all nine exploits are the result of a host component leaked into the container. Here, we apply a user namespace container defense to prevent the adversary from leveraging leaked host components and demonstrate that the defense stops seven of the nine container escape exploits.

42 ENGINEERING↗

Prototype Design of Global Common Module for ATLAS Experiment’s Phase-II Upgrade

A new Global Trigger subsystem will be installed in the Level-0 Trigger as part of HL-LHC Upgrade of ATLAS during the upcoming Long-Shutdown 3. It will feature new and improved trigger hardware and algorithms, and an increased maximum output rate of 1 MHz. The Global Trigger will run offline-like trigger algorithms on full-granularity data, gathered from several sub-detectors and trigger-processing subsystems. A single Global Common Module (GCM) hardware is implemented across the Global Trigger system to be used as Multiplexer Processor, Global Event Processor and CTP Interface (gCTPi). This common hardware platform method will minimize the complexity of the firmware and simplify the system design and long-term maintenance. The GCM prototype is an ATCA front form factor board with two Xilinx Virtex UltraScale+ FPGA VU13P and one ZYNQ UltraScale+ FPGA ZU19EG and seventeen 25.78125 Gb/s FireFly duplex optical modules on it. The total power consumption of this board must be less than 350 W, and the temperature of the optical modules should be less than 70 °C in the worst case. The VU13Ps serve as algorithms processor nodes such as MUX, GEP and gCTPi, and the ZU19EG with Peta Linux OS running on it, is used as Command/Control/Readout Unit to configure and monitor the board and communicate with the ATLAS Detector Control System (DCS). The development of an ATCA blade with three large FPGAs and about 200 optical links running at 25Gb/s is a very challenging task, and the successful test results have demonstrated this GCM prototype as an advancement of state-of-the-art electronics module design in HEP experiments. This paper presents the hardware design considerations, functionalities, and performance test results of this GCM prototype.

47 OTHER INSTRUMENTATION↗

Real-Time Ethernet Interface for NSTX-U’s Thomson Scattering Diagnostic (2023)

Here, the multipoint Thomson scattering (MPTS) diagnostic system at the National Spherical Torus Experiment Upgrade (NSTX-U) facility is undergoing an upgrade to operate in real-time and interface with the plasma control system (PCS) for NSTX-U. Previous prototyping efforts have shown that spectral analysis and rapid calculations of electron temperature and density are possible on a real-time Linux machine when using up to a 100-Hz laser pulse repetition rate. A remaining challenge was transferring the real-time data to NSTX-U’s PCS, which utilizes the front panel data port (FPDP) protocol. The original proposed method was to convert the real-time data into analog values, but a new solution was developed to keep the output format digital by using an Ethernet controller with a field-programmable gate array (FPGA). This article focuses on a new input module that has been developed to accept incoming user datagram protocol (UDP) packets sent over Ethernet, convert into FPDP format, and integrate into the existing data stream under NSTX-U’s real-time framework.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Can Applications Recover from fsync Failures?

We analyze how file systems and modern data-intensive applications react to fsync failures. First, we characterize how three Linux file systems (ext4, XFS, Btrfs) behave in the presence of failures. We find commonalities across file systems (pages are always marked clean, certain block writes always lead to unavailability) as well as differences (page content and failure reporting is varied). Next, we study how five widely used applications (PostgreSQL, LMDB, LevelDB, SQLite, Redis) handle fsync failures. Our findings show that although applications use many failure-handling strategies, none are sufficient: fsync failures can cause catastrophic outcomes such as data loss and corruption. Our findings have strong implications for the design of file systems and applications that intend to provide strong durability guarantees.

Computer Science↗

Preliminary Study on Fine-Grained Power and Energy Measurements on Grace Hopper GH200 with Open-Source Performance Tools

The increasing adoption of tightly integrated, heterogeneous architectures, combined with the slowdown of Moore’s law, has made application power and energy-driven optimizations critical to efficiently use high-performance computing systems. This paper introduces a newly developed open-source toolkit that seamlessly integrates the Linux real-time hardware monitoring program hwmon with the Performance Application Programming Interface and the Score-P performance measurement system, thereby enabling fine-grained power and energy measurements for high-performance computing applications. Our primary target platform is the Wombat test bed, which is a system based on the NVIDIA GH200 superchip. The toolkit can capture transient power peaks with high temporal resolution (50 ms) and, thanks to Score-P integration, can map power metrics to specific code regions, thereby providing actionable information on power-intensive operations and inefficiencies. The toolkit also provides a holistic view of both the power and the energy consumption of the entire GH200 superchip by covering all major components: the Grace CPU, the Hopper GPU, and the I/O subsystem. Experiments that use Locally Self-consistent Multiple Scattering, which is an application for first-principles calculations of materials developed at Oak Ridge National Laboratory, have demonstrated the tool’s ability to identify transient power spikes and uncover opportunities for energy-aware optimizations. Additionally, we introduce a Python-based utility for converting Open Trace Format 2 traces to Parquet format, thus enabling advanced data analysis for numerical integration methods applied to power data for accurate energy profiling.

Hernandez Mendoza, Oscar [ORNL] (ORCID:00000002538↗

Flexible and Effective Object Tiering for Heterogeneous Memory Systems

Computing platforms that package multiple types of memory, each with their own performance characteristics, are quickly becoming mainstream. To operate efficiently, heterogeneous memory architectures require new data management solutions that are able to match the needs of each application with an appropriate type of memory. As the primary generators of memory usage, applications create a great deal of information that can be useful for guiding memory management, but the community still lacks tools to collect, organize, and leverage this information effectively. To address this gap, this work introduces a novel software framework that collects and analyzes object-level information to guide memory tiering. The framework includes tools to monitor the capacity and usage of individual data objects, routines that aggregate and convert this information into tier recommendations for the host platform, and mechanisms to enforce these recommendations according to user-selected policies. Moreover, the developed tools and techniques are fully automatic, work on standard Linux systems, and do not require modification or recompilation of existing software. Using this framework, this study evaluates and compares the impact of a variety of design choices for memory tiering, including different policies for prioritizing objects for the fast memory tier as well as the frequency and timing of migration events. In conclusion, the results, collected on a modern Intel platform with conventional DDR4 SDRAM as well as Intel Optane NVRAM, show that guiding data tiering with object-level information can enable significant performance and efficiency benefits compared with standard hardware- and software-directed data-tiering strategies for a diverse set of memory-intensive workloads.

97 MATHEMATICS AND COMPUTING↗

DEDUPKV: A Space-Efficient and High-Performance Key-Value Store via Fine-Grained Deduplication

Log-Structured Merge Tree (LSM-tree) based key-value stores excel in write-intensive environments but suffer from data duplication, consuming up to 49% of storage space in LSM-tree-based key-value store deployments. Traditional solutions like compression and coarse-grained file system-level deduplication introduce overhead or have limited effectiveness. In this study, we propose DedupKV, a fine-grained deduplication framework tailored for LSM-tree, maximizing data reduction efficiency while minimizing write stalls and read overheads. DedupKV features three key innovations: (1) FLUSH-integrated inline deduplication, which removes duplicates during memory-to-storage writes; (2) WAL file-based offline deduplication, repurposing write-ahead logs to avoid double writes; and (3) elastic execution, dynamically balancing inline and offline deduplication based on memory pressure and workload intensity. Additionally, dynamic granularity management reduces deduplication metadata overhead. We implemented these four ideas in RocksDB for the first time and conducted experiments in a Linux environment. Our evaluation shows that WAL file-based offline deduplication and DedupKV outperform BlobDB by 33% and 23%, respectively, in write-heavy workloads, while reducing write amplification by 1.2 ×, 2 ×, and 1.6 × for real KV datasets.

Jamil, Safdar [Sogang University]↗

A software package for plasma facing component analysis and design: the Heat flux Engineering Analysis Toolkit (HEAT)

The engineering limits of plasma facing components (PFCs) constrain the allowable operational space of tokamaks. Poorly managed heat fluxes that push the PFCs beyond their limits not only degrade core plasma performance via elevated impurities, but can also result in PFC failure due to thermal stresses or melting. Simple axisymmetric assumptions fail to capture the complex interaction between 3D PFC geometry and 2D or 3D plasmas. This results in fusion systems that must either operate with increased risk or reduce PFC loads, potentially through lower core plasma performance, to maintain a nominal safety factor. High precision 3D heat flux predictions are necessary to accurately ascertain the state of a PFC given the evolution of the magnetic equilibrium. A new code, the Heat flux Engineering Analysis Toolkit (HEAT), has been developed to provide high precision 3D predictions and analysis for PFCs. HEAT couples many otherwise disparate computational tools together into a single open source python package. Magnetic equilibrium, engineering CAD, finite volume solvers, scrape off layer plasma physics, visualization, high performace computing, and more, are connected in a single web-based user interface. Linux users may use HEAT without any software prerequisites via an appImage. This manuscript introduces HEAT, discusses the software architecture, presents first HEAT results, and outlines physics modules in development.

divertor physics↗

Parameter Optimization Toolbox for NS-3 network optimization, NS-3 Parameter Optimization Framework [SWR-18-60]

This simulation-based parameter optimization framework is proposed to tune parameters of different types of communication networks using ns-3 to achieve the optimal network performance. It consists of three main components: an ns-3 packet reporting module; a sampler running simulations with all possible parameter sets for the input parameter variables by using a parallel executor at each generation; and a hybrid optimization algorithm for tuning configurable parameters of hybrid designs and application parameter variables. The proposed hybrid metaheuristic optimization algorithm combines an evolutionary algorithm with a gradient descent function to quickly achieve an approximate globally optimum solution. This software is designed to be used in a multi-core processing Linux environment and run over a long duration of time. The execution time varies depending mainly upon the nature of the ns-3 configuration being simulated. This software includes a custom ns-3 QoS measurement application which must be included with the ns-3 source code during installation of the software.

Hasandka, Adarsh↗

Logic in Memory Emulator

Logic in Memory Emulator (LiME) is a hardware/software tool specially designed for memory system evaluation and experiment. Emerging memories display a wide range of bandwidths, latencies, and capacities, making it challenging for the computer architect to navigate the design space of potential memory configurations, and for the application developer to assess performance implications of using such memories. With the LiME framework, architectural ideas can be prototyped in great detail yet with sufficient performance to support realistic evaluation on long running applications. LiME consists of two fundamental components: 1) the hardware and OS infrastructure for the emulator, and 2) a suite of benchmark applications to assist in characterizing the performance of current and future computer architectures. Some of the applications have been collected from other open source projects. Uses: Logging, replay and analysis of an application's memory behavior Evaluate impact of emerging memory technology on application performance. Emulate complex memory interactions in whole applications orders of magnitude faster than software simulation. Emulate acceleration hardware co-located with the memory subsystem. Features: Capture and log external memory accesses to a separate off-chip memory device without affecting application execution. Memory traces include the address, length, timestamp, and optionally the data for each transaction. Captured trace data can be saved to an SD card for off-line analysis. Configure a wide range of memory latencies in sub-nanosecond increments that encompass highbandwidth and storage class memories. Specify regions of interest (ROI) in applications to reduce the amount of trace data captured for analysis. Currently supports execution on Xilinx Zynq SoC which integrates an ARM processor with FPGA logic on a single device. Applications can be run under Linux or in bare metal mode on the ARM cores.

Jain, AbhishekK↗

Greggd

greg(g)d - Global runtime for eBPF-enabled gathering (w/ gumption) daemon Recently the linux kernel has added support for low-level kernel monitoring and profiling through a in-kernel virtual machine. The tooling around these new features (the extended Berkley Packet Filter or eBPF for short) is not mature and is difficult to use. Benefits from eBPF are especially hard to realize while trying to do large scale deployments and integrate with existing metric analysis stacks. A tool was needed to enable loading and collecting data from eBPF programs on large scale HPC systems. Given the problems above it was obvious we needed some wrapper program to compile, load, and collect data from eBPF programs running in the kernel. This tool needed to be lightweight without a heavy set of dependencies, relatively stable between different kernel versions, and integrate nicely with existing widely used metric collection tools. We wrote a program that wraps the eBPF tooling and sends data to our metric gathering tool. eBPF programs are either compiled using the host compiler stack, or loaded in the kernel directly from an object file. These programs are then attached to the system calls that we want to profile. Whenever these system calls are run, the eBPF program collects information of interest and writes that to memory. Our wrapper program polls these memory locations, reads and formats the data, then sends the information to a local unix socket. Our other monitoring tools are configured to read from that socket and send it to the rest of our metric monitoring stack for analysis.

Voss, Joseph [Oak Ridge National Lab. (ORNL), Oak ↗

Advanced Tri-lab Software Environment (ATSE)

The Advanced Tri-lab Software Environment (ATSE) is an effort to build an open, modular, extensible, community-engaged, and vendor-adaptable software ecosystem that enables the prototyping of new technologies for improving the ASC computing environment. The initial target for ATSE is to accelerate the maturity of the Arm ecosystem for supporting ASC computing and the high-performance computing community more broadly. ATSE provides an integrated and optimized software stack that includes: 1) Application development environment and libraries including compilers, math libraries, tools, MPI, and OpenMP, 2) Low-level system software including optimized Linux, network stack, file systems, containers, and virtual machines, 3) Job scheduling and management including workload manager, application launcher, and user tools, and 4) System administration and management tools supporting booting, monitoring, and operating system image management. SAND2020-12377 M Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Younge, Andrew J.↗

Network Architecture Verification & Validation Tool

The NAVV Tool is an automation of Linux commands run Zeek IDS software on a packet capture to create a Microsoft Excel spreadsheet table breaking down network traffic observed. The tool automates the Zeek software analysis, the collation of logs, and then the dissection of the Conn.log and DNS.logs to create a summary table within a Excel. This spreadsheet can then be updated with network segments using CIDR formatting and labels along with inventory information including name and IP address. Using the tool again will integrate these label and color coding into the existing analysis table to aid in conducting an evaluation of the network traffic.

Nichols, DonovanW↗