Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “task-based”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Performance of Heterogeneous Algorithm Scheduling in CMSSW

The CMS experiment started to utilize Graphics Processing Units (GPU) to accelerate the online reconstruction and event selection running on its High Level Trigger (HLT) farm in the 2022 data taking period. The projections of the HLT farm to the High-Luminosity LHC foresee a significant use of compute accelerators in the LHC Run 4 and onwards in order to keep the cost, size, and power budget of the farm under control. This direction of leveraging compute accelerators has synergies with the increasing use of HPC resources in HEP computing, as HPC machines are employing more and more compute accelerators that are predominantly GPUs today. In this work we review the features developed for the CMS data processing framework, CMSSW, to support the effective utilization of both compute accelerators and many-core CPUs within a highly concurrent task-based framework. We measure the impact of various design choices for the scheduling of heterogeneous algorithms on the event processing throughput, using the Run-3 HLT application as a realistic use case.

Bocci, Andrea↗

Implementation of compound refractive lenses for large field-of-view x-ray phase-contrast imaging during hypervelocity impact experiments

Synchrotron x-ray phase-contrast imaging (XPCI) offers time-resolved visualization of dynamic compression phenomena, but its intrinsically small field-of-view (FOV) limits the time that key features remain in frame. A novel approach to enlarge the FOV is achieved by positioning a two-dimensional parabolic compound refractive lens (CRL) upstream of the sample to deliberately defocus the white beam. Ray-tracing simulations and XPCI measurements show that this CRL configuration can expand the beam by ∼50% vertically and ∼15% horizontally based on the full width at half-maximum of the beam. Implementing the CRL, however, attenuates the photon flux and lowers signal-to-noise ratio (SNR). Task-based analysis using a calibration grid (30 μm dots) showed that both setups fail to consistently meet the Rose criterion (SNR ≥ 5) for features of this size in single-bunch imaging. Extrapolating the measured SNR Rose values suggests that the minimum consistently detectable feature lies closer to 30–40 μm for the standard XPCI setup and above 40 μm for CRL-XPCI. Despite this limitation, the CRL configuration nearly doubles the illuminated area, enabling simultaneous tracking of front and rear observations of boron carbide targets subjected to rod and sphere impacts at 1.0–2.6 km/s. Image tracking algorithms and photonic Doppler velocimetry were used to measure penetration and rear-surface velocity histories. Together, these measurements capture crack fronts, penetration, and material breakout, offering new benchmark data for validating high-strain-rate constitutive models of ceramic materials.

Ceramic materials↗

Automated segmentation of soft X-ray tomography: Native cellular structure with submicron resolution at high-throughput for whole-cell quantitative imaging in yeast

Soft X-ray tomography (SXT) is an invaluable tool for quantitatively analyzing cellular structures at suboptical isotropic resolution. However, it has traditionally depended on manual segmentation, limiting its scalability for large datasets. Here, we leverage a deep learning-based autosegmentation pipeline to segment and label cellular structures in hundreds of cells across three Saccharomyces cerevisiae strains. This task-based pipeline uses manual iterative refinement to improve segmentation accuracy for key structures, including the cell body, nucleus, vacuole, and lipid droplets, enabling high-throughput and precise phenotypic analysis. Using this approach, we quantitatively compared the three-dimensional (3D) whole-cell morphometric characteristics of wild-type, VPH1-GFP, and vac14 strains, uncovering detailed strain-specific cell and organelle size and shape variations. We show the utility of SXT data for precise 3D curvature analysis of entire organelles and cells and detection of fine morphological features using surface meshes. Our approach facilitates comparative analyses with high spatial precision and statistical throughput, uncovering subtle morphological features at the single-cell and population level. This workflow significantly enhances our ability to characterize cell anatomy and supports scalable studies on the mesoscale, with applications in investigating cellular architecture, organelle biology, and genetic research across diverse biological contexts.

Chen, Jianhua [Lawrence Berkeley National Laborato↗

IRIS-DMEM: Efficient Memory Management for Heterogeneous Computing

This paper proposes an efficient data memory management approach for the Intelligent RuntIme System (IRIS) heterogeneous computing framework along with new data transfer policies. IRIS provides a task-based programming model for extreme heterogeneous computing (e.g., CPU, GPU, DSP, FPGA) with support for today's most important programming languages (e.g., OpenMP, OpenCL, CUDA, HIP, OpenACC). However, the IRIS framework either forces the programmer to introduce data transfer commands for each task or relies on suboptimal memory management for automatic and transparent data transfers. The work described here extends IRIS with novel heterogeneous memory handling and introduces novel data transfer policies by employing the Distributed data MEMory handler (DMEM) for efficient and optimal movement of data among the various computing resources. The proposed approach achieves performance gains of up to 7× for tiled LU factorization and tiled DGEMM (i.e., matrix multiplication) benchmarks. Moreover, this approach also reduces data transfers by up to 71% when compared to previous IRIS heterogeneous memory management handlers. This work compares the performance results of the IRIS framework's novel DMEM with the StarPU runtime and MAGMA math library for GPUs. Experiments show a performance gain of up to 1.95× over StarPU and 2.1× over MAGMA.

Miniskar, Narasinga Rao↗

IRIS-MEMFLOW: Data Flow-Enabled Portable Memory Orchestration in IRIS Runtime for Diverse Heterogeneity

Task-based programming models and execution paradigms provide a means to decompose a computation by expressing it as a graph in which each node represents a specific computation operating on memory objects and the edges define the dependencies in the execution flow. In this execution model, independent nodes in the graph can be executed concurrently in different computing devices, making it suitable for heterogeneous systems in which computing devices with different architectures coexist. However, careful memory orchestration across heterogeneous devices is needed because copies of the same memory object may reside in multiple devices during execution. Manually ensuring such an orchestration is quite challenging. Not only must an application developer guard against race conditions, but they must also optimize data movement between the host and devices because unnecessary data movement significantly impacts performance. To mitigate these challenges, we enhance the IRIS heterogeneous runtime and introduce IRIS-MEMFLOW–a data flow–enabled portable memory abstraction for seamlessly orchestrating memory in diverse heterogeneous computing environments. By using data-flow analysis, IRIS-MEMFLOW guards against race conditions while multiple heterogeneous devices access memory objects. IRIS-MEMFLOW also optimizes data movement between the host and devices without manual intervention. As a result, IRIS provides improved programming productivity, performance, and portability for multidevice heterogeneous executions in high-performance computing and cloud systems that run diverse architectures from different vendors. The efficacy of IRIS-MEMFLOW is evaluated through experiments that show its capability in terms of programming productivity, multidevice heterogeneity, portability, and low overhead versus the state of the art.

Monil, M. A. H. [ORNL] (ORCID:0000000334194037)↗

Streaming Statistics

In the context of a larger effort for in situ data analytics, there is a need to calculate basic statistics metrics (e.g., count, mean, median) online as new data points become available. Originally, the code for such online, or in other words streaming, statistics was part of the TALASS (Topological Analysis of Large- Scale Simulations) library. We isolated the relevant code and created a standalone library from it called Streaming Statistics. We also added an ability to serialize and deserialize the statistics objects so that the library can be used in distributed, task-based processing. To use the Streaming Statistics library, the user chooses a statistic, constructs an object for it, and then "adds" values to it, which means the statistic is augmented.

Shudler, Sergei↗

pnnl/mcl-runtime

The Minos Computing Library (MCL) is a task-based programming language and runtime for extremely heterogeneous systems. MCL facilitates writing program for heterogeneous devices and porting applications across different systems and devices. MCL supports asynchronous execution of computing tasks on all available heterogeneous devices, including GPUs, FGPAs, fixed-point accelerators, and AI accelerators. MCL provides a high-level programming interface and automatically and autonomously performs resource management, load balancing, and locality-aware scheduling.

Central, PNNL Developer↗

An Investigation of using FleCSI for Monte Carlo Radiation Transport

This document details the attempt to use FleCSI to provide MPI parallelization and domain decomposition for a Monte Carlo radiation transport application. FleCSI is a framework designed to support multi-physics application development with a focus on task-based parallelism and domain decomposition [1]. FleCSI also supports performance portability via a wrapper around a back-end performance portability layer. The goal of this work is to use FleCSI for domain decomposition and MPI parallelization inside of a Monte Carlo radiation transport application and document any pain points and shortcomings. The following section introduce the nomenclature used in FleCSI and its potential benefit for physics code developers, detail the code used to explore using FleCSI in a Monte Carlo radiation transport solver, and list the issues and concerns discovered during the work.

36 MATERIALS SCIENCE↗

Heterogeneous Computing

To leverage the increasing heterogeneity in modern computing resources, Geant4 incorporates advanced software tools and a task-based framework (G4Tasking) that enables efficient parallelism at event, sub-event, and track levels. Ongoing R&D efforts focus on integrating GPUs into high-energy physics (HEP) simulations, including optical photon simulation with Opticks/NVIDIA OptiX, offloading electromagnetic particle transport using G4HepEM/AdePT and Celeritas, and employing advanced surface-based geometry models such as VecGeom2.0 and ORANGE. As Geant4 continues evolving toward high-performance computing (HPC) and heterogeneous architectures, it remains a key tool for large-scale simulations in HEP and beyond.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Parallel Programming in MCNP6

Monte Carlo N-Particle (MCNP)1 is a general-purpose Monte Carlo particle transport code developed by Los Alamos National Laboratory (LANL). To efficiently handle long simulations, MCNP version 6 (MCNP6) supports parallel execution using two primary programming models: • Shared-memory task-based threading using OpenMP (Open Multi-Processing), and • Distributed-memory calculations using MPI (Message Passing Interface). The OpenMP and MPI programming models enable MCNP6 to scale from desktop systems to high-performance computing (HPC) clusters, allowing users to run MCNP in one of three parallel modes: • OpenMP-only, • MPI-only, and • Hybrid (MPI + OpenMP). The choice of parallelization mode depends on the underlying computer architecture and the characteristics of the simulation problem.

97 MATHEMATICS AND COMPUTING↗

Sound-localization-related activation and functional connectivity of dorsal auditory pathway in relation to demographic, cognitive, and behavioral characteristics in age-related hearing loss

Background Patients with age-related hearing loss (ARHL) often struggle with tracking and locating sound sources, but the neural signature associated with these impairments remains unclear. Materials and methods Using a passive listening task with stimuli from five different horizontal directions in functional magnetic resonance imaging, we defined functional regions of interest (ROIs) of the auditory “where” pathway based on the data of previous literatures and young normal hearing listeners ( n = 20). Then, we investigated associations of the demographic, cognitive, and behavioral features of sound localization with task-based activation and connectivity of the ROIs in ARHL patients ( n = 22). Results We found that the increased high-level region activation, such as the premotor cortex and inferior parietal lobule, was associated with increased localization accuracy and cognitive function. Moreover, increased connectivity between the left planum temporale and left superior frontal gyrus was associated with increased localization accuracy in ARHL. Increased connectivity between right primary auditory cortex and right middle temporal gyrus, right premotor cortex and left anterior cingulate cortex, and right planum temporale and left lingual gyrus in ARHL was associated with decreased localization accuracy. Among the ARHL patients, the task-dependent brain activation and connectivity of certain ROIs were associated with education, hearing loss duration, and cognitive function. Conclusion Consistent with the sensory deprivation hypothesis, in ARHL, sound source identification, which requires advanced processing in the high-level cortex, is impaired, whereas the right–left discrimination, which relies on the primary sensory cortex, is compensated with a tendency to recruit more resources concerning cognition and attention to the auditory sensory cortex. Overall, this study expanded our understanding of the neural mechanisms contributing to sound localization deficits associated with ARHL and may serve as a potential imaging biomarker for investigating and predicting anomalous sound localization.

Wu, Junzhi↗

Configuration control of redundant manipulators - Theory and implementation

A simple approach for controlling the manipulator configuration over the entire motion is presented, based on augmentation of the manipulator forward kinematics. User-defined kinematic functions and the end-effector Cartesian coordinates are combined to form a set of task-related configuration variables as generalized coordinates for the manipulator. A task-based adaptive scheme is then utilized to control the configuration variables and achieve tracking of the desired reference trajectories. This achieves the desired end-effector motion while utilizing redundancy to achieve any additional task. Simulation results for a direct-drive two-link arm are given to illustrate the proposed control scheme. The scheme has also been implemented for real-time control of three links of a PUMA 560 industrial robot. The simulation and experimental results validate the configuration control scheme and demonstrate its capabilities for performing various realistic tasks.

Seraji, Homayoun↗

Method and apparatus for configuration control of redundant robots

A method and apparatus to control a robot or manipulator configuration over the entire motion based on augmentation of the manipulator forward kinematics is disclosed. A set of kinematic functions is defined in Cartesian or joint space to reflect the desirable configuration that will be achieved in addition to the specified end-effector motion. The user-defined kinematic functions and the end-effector Cartesian coordinates are combined to form a set of task-related configuration variables as generalized coordinates for the manipulator. A task-based adaptive scheme is then utilized to directly control the configuration variables so as to achieve tracking of some desired reference trajectories throughout the robot motion. This accomplishes the basic task of desired end-effector motion, while utilizing the redundancy to achieve any additional task through the desired time variation of the kinematic functions. The present invention can also be used for optimization of any kinematic objective function, or for satisfaction of a set of kinematic inequality constraints, as in an obstacle avoidance problem. In contrast to pseudoinverse-based methods, the configuration control scheme ensures cyclic motion of the manipulator, which is an essential requirement for repetitive operations. The control law is simple and computationally very fast, and does not require either the complex manipulator dynamic model or the complicated inverse kinematic transformation. The configuration control scheme can alternatively be implemented in joint space.

Seraji, Homayoun↗

Scheduling Operations for Massive Heterogeneous Clusters

High-performance computing (HPC) programming has become increasingly difficult with the advent of hybrid supercomputers consisting of multicore CPUs and accelerator boards such as the GPU. Manual tuning of software to achieve high performance on this type of machine has been performed by programmers. This is needlessly difficult and prone to being invalidated by new hardware, new software, or changes in the underlying code. A system was developed for task-based representation of programs, which when coupled with a scheduler and runtime system, allows for many benefits, including higher performance and utilization of computational resources, easier programming and porting, and adaptations of code during runtime. The system consists of a method of representing computer algorithms as a series of data-dependent tasks. The series forms a graph, which can be scheduled for execution on many nodes of a supercomputer efficiently by a computer algorithm. The schedule is executed by a dispatch component, which is tailored to understand all of the hardware types that may be available within the system. The scheduler is informed by a cluster mapping tool, which generates a topology of available resources and their strengths and communication costs. Software is decoupled from its hardware, which aids in porting to future architectures. A computer algorithm schedules all operations, which for systems of high complexity (i.e., most NASA codes), cannot be performed optimally by a human. The system aids in reducing repetitive code, such as communication code, and aids in the reduction of redundant code across projects. It adds new features to code automatically, such as recovering from a lost node or the ability to modify the code while running. In this project, the innovators at the time of this reporting intend to develop two distinct technologies that build upon each other and both of which serve as building blocks for more efficient HPC usage. First is the scheduling and dynamic execution framework, and the second is scalable linear algebra libraries that are built directly on the former.

Humphrey, John↗

Neural Predictors of Visuomotor Adaptation Rate and Multi-Day Savings

Recent studies of sensorimotor adaptation have found that individual differences in task-based functional brain activation are associated with the rate of adaptation and savings at subsequent sessions. However, few studies to date have investigated offline neural predictors of adaptation and multi-day savings. In the present study, we explore whether individual differences in the rate of visuomotor adaptation and multi-day savings are associated with differences in resting state functional connectivity and gray matter volume. Thirty-four participants performed a manual adaptation task during two separate test sessions, on average 9 days apart. We found that resting state functional connectivity strength between sensorimotor, anterior cingulate, and temporoparietal areas of the brain was a significant predictor of adaptation rate during the early, cognitive phase of practice. In contrast, default mode network functional connectivity strength was found to predict late adaptation rate and savings on day two, which suggests that these behaviors may rely on overlapping processes. We also found that gray matter volume in temporoparietal and occipital regions was a significant predictor of early learning, whereas gray matter volume in superior posterior regions of the cerebellum was a significant predictor of late adaptation. The results from this study suggest that offline neural predictors of early adaptation facilitate the cognitive mechanisms of sensorimotor adaptation, with support from by the involvement of temporoparietal and cingulate networks. In contrast, the neural predictors of late adaptation and savings, including the default mode network and the cerebellum, likely support the storage and modification of newly acquired sensorimotor representations. These findings provide novel insights into the neural processes associated with individual differences in sensorimotor adaptation.

Cassady, Kaitlin↗

MEXEC: An Onboard Integrated Planning and Execution Approach for Spacecraft Commanding

The traditional form of spacecraft commanding is with sequences that specify when commands should execute based on a schedule generated on the ground. Some sequences have control logic and event driven responses to increase flexibility, but it is limited. An approach to increase autonomy is to use goal-based planning and commanding. Using this paradigm, intention and behavior is modeled on board the spacecraft. In this paper we describe MEXEC (Multi-mission EXECutive), a multi-mission, task-based, onboard planning and execution software designed specifically to be used as flight software. As a path to infusion for future flight projects, we describe two experiments performed on the ASTERIA CubeSat and testbed that demonstrate that MEXEC can be integrated and used for spacecraft operations and increase robustness and science return compared to the standard sequences that were being used.

Campuzano, Brian↗

Validation of Fitness for Duty Standards Using Pre- and Post-Flight Capsule Egress and Suited Functional Performance Tasks in Simulated Reduced Gravity

The transition between gravity environments will involve one of the most complex, high-risk phases of any mission. For example, the reduced functional capacity caused by physiological deconditioning adaptations in microgravity coupled with the stressors of re-entry into partial gravity environments will increase risks to crew, even with rigorous adherence to inflight countermeasures. Quantification of the astronauts’ post-landing functional performance is necessary to design ConOps for exploration missions. Specifically, these two high-risk scenarios may be required to be performed soon after gravity transitions: • Nominal and/or emergency unassisted capsule egress task after return to Earth • Planetary extravehicular activity (EVA) soon after landing on Mars or the Moon This study is broken down into two phases. Phase 1: A feasibility study that will assess the overall feasibility and demonstrate the capability to do these tasks shortly after landing. Phase 2: the full Egress Fitness study, which is part of the CIPHER complement. This study uses a task-based approach to characterize functional performance of these high-risk scenarios in long-duration ISS crewmembers before flight and shortly after return to Earth. Prior to any testing, each astronaut subject completes a suit fit check to ensure adequate sizing and a mobility assessment to confirm completion of the EVA tasks. Pilot Egress Fitness pre-flight and post-flight testing includes an Earth-based emergency egress out of a functional capsule mockup and a short Mars gravity EVA simulation including suit donning, hatch egress, ladder descent, task board cable operations, baggage transfer over sand/rocky regolith, alignment with a rear entry port, and suit egress. Pre-flight testing will occur at any time point prior to flight. Post-flight assessment is much more critical with the capsule egress test that will occur 1–4 hours after landing and the planetary EVA approximately 18–36 hours after landing. The CIPHER study will incorporate additional pre-flight sessions, longer EVA tasks that include traverse and geology sampling, and post-flight sessions on R+1, 4, and 8 to characterize the timeframe of recovery. Data collected for both tasks include task completion time, photo, and video. The EVA portion also includes collection of metabolic and heart rate. Pilot Egress Fitness study has completed both baseline and post-flight testing on four astronaut subjects, with all four able to complete post-flight EVA testing and three able to complete postflight capsule egress testing. Preliminary results observe individual physiological variations, which were to be expected but also suggest the need to carefully track the timeline from undock to landing to testing. Furthermore, some task performance instructions and equipment may need to be adjusted to ensure results are primarily physiological. These and other lessons learned will be addressed with minor protocol changes in the full CIPHER Egress Fitness study.

J R Norcross↗