Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Runtime Scheduling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Asynchronous distributed-memory task-parallel algorithm for compressible flows on unstructured 3D Eulerian grids

Here, we discuss the implementation of a finite element method, used to numerically solve the Euler equations of compressible flows, using an asynchronous runtime system (RTS). The algorithm is implemented for distributed-memory machines, using stationary unstructured 3D meshes, combining data-, and task-parallelism on top of the Charm++ RTS. Charm++’s execution model is asynchronous by default, allowing arbitrary overlap of computation and communication. Task-parallelism allows scheduling parts of an algorithm independently of, or dependent on, each other. Built-in automatic load balancing enables continuous redistribution of computational load by migration of work units based on real-time CPU load measurement. The RTS also features automatic checkpointing, fault tolerance, resilience against hardware failure, and supports power-, and energy-aware computation. We demonstrate scalability up to 25 x 10 9 cells at $\mathscr{O}$10 4 compute cores and the benefits of automatic load balancing for irregular workloads. The full source code with documentation is available at https://quinoacomputing.org.

42 ENGINEERING↗

Concurrent Runtime Verification of Data Rich Events

This paper presents the open source runtime verification tool MESA (MEssage-based System Analysis), implemented in Scala, which supports concurrent monitors using the Actor model. Furthermore, the tool supports indexing (slicing) on the data values occurring in data-carrying events, for each individual monitor. The tool is generic in the sense that any monitoring system can be used for creating monitors. In this paper, we use the internal Scala DSL Daut for programming such in data parameterized state machines and temporal logic. To illustrate MESA/Daut, we present a case study that monitors flights from live U.S. airspace data streams, verifying that they conform to planned routes. With base in the case study, we then perform an extensive empirical study of the potential benefits from monitoring slices of a single property in concurrently executing actors. Due to the overhead of scheduling “small” actors (one for each slice or a small number of slices), it is not obvious that concurrent execution of such is beneficial. However, as a main result, we demonstrate that concurrent monitoring of slices to handle data-carrying events can provide considerable speed gains.

finite state machines↗

Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix Multiplication.

There is a growing interest in custom spatial accelerators for machine learning applications. These accelerators employ a spatial array of processing elements (PEs) interacting via custom buffer hierarchies and networks-on-chip. The efficiency of these accelerators comes from employing optimized dataflow (i.e., spatial/temporal partitioning of data across the PEs and fine-grained scheduling) strategies to optimize data reuse. The focus of this work is to evaluate these accelerator architectures using a tiled general matrix-matrix multiplication (GEMM) kernel. To do so, we develop a framework that finds optimized mappings (dataflow and tile sizes) for a tiled GEMM for a given spatial accelerator and workload combination, leveraging an analytical cost model for runtime and energy. Finally, our evaluations over five spatial accelerators demonstrate that the tiled GEMM mappings systematically generated by our framework achieve high performance on various GEMM workloads and accelerators.

42 ENGINEERING↗

Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix Multiplication

There is a growing interest in custom spatial accelerators for machine learning applications. These accelerators employ a spatial array of processing elements (PEs) interacting via custom buffer hierarchies and networks-on-chip. The efficiency of these accelerators comes from employing optimized dataflow (i.e., spatial/temporal partitioning of data across the PEs and fine-grained scheduling) strategies to optimize data reuse. The focus of this work is to evaluate these accelerator architectures using a tiled general matrix-matrix multiplication (GEMM) kernel. To do so, we develop a framework that finds optimized mappings (dataflow and tile sizes) for a tiled GEMM for a given spatial accelerator and workload combination, leveraging an analytical cost model for runtime and energy. Our evaluations over five spatial accelerators demonstrate that the tiled GEMM mappings systematically generated by our framework achieve high performance on various GEMM workloads and accelerators.

43 PARTICLE ACCELERATORS↗

pnnl/mcl-runtime

The Minos Computing Library (MCL) is a task-based programming language and runtime for extremely heterogeneous systems. MCL facilitates writing program for heterogeneous devices and porting applications across different systems and devices. MCL supports asynchronous execution of computing tasks on all available heterogeneous devices, including GPUs, FGPAs, fixed-point accelerators, and AI accelerators. MCL provides a high-level programming interface and automatically and autonomously performs resource management, load balancing, and locality-aware scheduling.

Central, PNNL Developer↗

Evapotranspiration-Based Irrigation Scheduling in Cool-Season Vegetables

Crop evapotranspiration (ETc) is strongly linked with photosynthetically active vegetation fraction (Fc). Estimation of ETc may support efficiency gains in irrigation water management, which in turn can mitigate nitrate leaching, promote water supply sustainability, and reduce energy costs associated with water pumping or transport. The University of California Cooperative Extension operates the CropManage (CM) model as a freely-available web-application for growers and consultants to support irrigation and nitrogen scheduling decisions. CM accounts for the rapid growth and typically brief cycle of cool-season vegetables, where Fc and crop coefficient (fraction of reference ET) can change daily during canopy development. Daily weather conditions are inherently accounted for by use of grass reference ETo data imported from the California Dept. Water Resources. Crop water requirement calculations are output in terms of irrigation system runtime. Empirical equations are used to estimate daily Fc time-series for a given crop type, primarily as a function of planting date and expected harvest. An applications programming interface (API) enables CM to import satellite-based Fc observations from NASA's Satellite Irrigation Management Support, which uses Landsat imagery to monitor about eight million irrigation acres statewide. The API is intended to provide a check on internal CM predictions of Fc and to facilitate expansion of the web-app to new crops and regions. A replicated irrigation trial was performed on cauliflower during spring/summer 2018 at the USDA Agricultural Research Station in Salinas, CA. The crop was established by sprinkler irrigation, and CropManage was then used to guide a series of drip irrigation treatments at 50%, 75%, 100%, and 150% of ETc replacement levels. Results will be presented with respect to water use efficiency, nitrogen use efficiency, biomass yield, and marketable yield. Additional findings will be presented for a celery trial harvested during autumn 2018.

Johnson, Lee↗

Runtime support and compilation methods for user-specified data distributions

This paper describes two new ideas by which an HPF compiler can deal with irregular computations effectively. The first mechanism invokes a user specified mapping procedure via a set of compiler directives. The directives allow use of program arrays to describe graph connectivity, spatial location of array elements, and computational load. The second mechanism is a simple conservative method that in many cases enables a compiler to recognize that it is possible to reuse previously computed information from inspectors (e.g. communication schedules, loop iteration partitions, information that associates off-processor data copies with on-processor buffer locations). We present performance results for these mechanisms from a Fortran 90D compiler implementation.

Ponnusamy, Ravi↗

Middleware and Web Services for the Collaborative Information Portal of NASA's Mars Exploration Rovers Mission

We describe the design and deployment of the middleware for the Collaborative Information Portal (CIP), a mission critical J2EE application developed for NASA's 2003 Mars Exploration Rover mission. CIP enabled mission personnel to access data and images sent back from Mars, staff and event schedules, broadcast messages and clocks displaying various Earth and Mars time zones. We developed the CIP middleware in less than two years time usins cutting-edge technologies, including EJBs, servlets, JDBC, JNDI and JMS. The middleware was designed as a collection of independent, hot-deployable web services, providing secure access to back end file systems and databases. Throughout the middleware we enabled crosscutting capabilities such as runtime service configuration, security, logging and remote monitoring. This paper presents our approach to mitigating the challenges we faced, concluding with a review of the lessons we learned from this project and noting what we'd do differently and why.

Sinderson, Elias↗

Execution-Based Model Checking of Interrupt-Based Systems

Execution-based model checking (EMC) is a verification technique based on executing a multi-threaded/multiprocess program repeatedly in a systematic manner in order to explore the different interleavings of the program. This is in contrast to traditional model checking, where a model of a system is analyzed Several execution-based model-checking tools exist at this point, such as for example Verisoft and Java PathFinder. The most common formal specification languages used by EMC tools are un- timed, either just assertions, or linear-time temporal logic (LTL). An alternative verification technique is Runtime Execution Monitoring (REM), which is based on monitor- ing the execution of a program, checking that the execution trace conforms to a requirement specification. The Temporal Rover and DBRover are such tools. They provide a very rich specification language, being an extension of LTL with real-time constraints and time-series. We show how execution-based model checking, combined with runtime execution monitoring, can be used for the verification of a large class of safety critical systems commonly known as interrupt-based systems. The proposed approach is novel in that: (i) it supports model checking of a large class of applications not practically verifiable using conventional EMC tools, (ii) it supports verification of LTL assertions extended with real-time and time-series constraints, and (iii) it supports the verification of custom schedulers.

Drusinsky, Doron↗

Hybrid PDES Simulation of HPC Networks Using Zombie Packets

Although high-fidelity network simulations have proven to be reliable and cost-effective tools to peer into architectural questions for high-performance computing (HPC) networks, they incur a high resource cost. The time spent in simulating a single millisecond of network traffic in the highest detail can take hours, even for static, well-behaved traffic patterns such as uniform random. Surrogate models offer a significant reduction in runtime, yet they cannot serve as complete replacements and should only be used when appropriate. Thus, there is a need for hybrid modeling, where high-fidelity simulation and surrogates run side-by-side. Here, we present a surrogate model for HPC networks in which: packets bypass the network, while the network state is left untouched, i.e., suspended. To bypass the network, we use historical data to estimate the arrival time at which every packet should be scheduled at; to suspend the network, all in-flight packets are scheduled to arrive at their destinations, and are kept in the system to awaken as zombies when switching back to high-fidelity. Speedup for a hybrid model is relative to the proportion of surrogate to high-fidelity. This light-weight surrogate obtained up to 76× speedup. Keeping the zombies in the network showed an increase in the accuracy of the high-fidelity simulation on restart when compared to restarting the network from an empty state.

HPC networks↗

Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis

HPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science, business, and other decision-making processes. However, understanding how ML jobs impact the operation of HPC datacenters, relative to generic jobs, remains desirable but understudied. In this work, we leverage long-term operational data, collected from a national-scale production HPC datacenter, and statistically compare how ML and generic jobs can impact the performance, failures, resource utilization, and energy consumption of HPC datacenters. Our study provides key insights, e.g., ML-related power usage causes GPU nodes to run into temperature limitations, median/mean runtime and failure rates are higher for ML jobs than for generic jobs, both ML and generic jobs exhibit highly variable arrival processes and resource demands, significant amounts of energy are spent on unsuccessfully terminating jobs, and concurrent jobs tend to terminate in the same state. We open-source our cleaned-up data traces on Zenodo (https://doi. org/10.5281/zenodo.13685426), and provide our analysis toolkit as software hosted on GitHub (https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization). This study offers multiple benefits for data center administrators, who can improve operational efficiency, and for researchers, who can further improve system designs, scheduling techniques, etc.

crossanalysis↗

CropManage Application for Vineyard Irrigation Decision-Support

CropManage is a free web-application developed by U.C. Cooperative Extension to support evapotranspiration based irrigation scheduling and nutrient management for major specialty crops. Prescribed phenology curves are used to develop daily estimates of canopy cover within a given field, based on days since planting (annual crops) or budbreak (trees, vines). These curves are modulated by a MaxCan parameter representing seasonal maximum canopy cover. Crop development observations can be used to adjust for such factors as weather anomalies or non-standard agronomic practice, as needed. Canopy cover is converted to crop coefficient and combined with reference evapotranspiration to derive daily water consumption. Guidance on crop water requirement is then conveyed to users in terms of system runtime issued on-demand for a given date, largely based on total evapotranspiration since last irrigation event. In this study, CropManage was adapted to vineyards by adding modules accounting for early-season soil moisture depletion and cover crop presence. A crop stress parameter was added to accommodate deficit irrigation practice, allowing the user to specify percentage departure from full water requirement along with start/stop dates. An initial verification exercise was performed on three winegrape vineyards located in California’s Central Coast (2020), North Coast (2020) and Central Valley (2019). Daily crop evapotranspiration was monitored by eddy-covariance fluxtowers. MaxCan was measured by ground and satellite observation. Stress regime was specified by grower practice where available, otherwise stress levels were inferred from applied water records. Mean absolute error and mean bias error of modeled cumulative evapotranspiration were computed with respect to the eddy covariance measurements collected throughout the growing season. Results indicate the modified CropManage water management module performs reasonably well for winegrape. Additional effort is planned to modify the nutrient module for vineyard use.

CropManage↗

Initial Evaluation of CropManage Decision-Support Model for Vineyard ET Estimation

The CropManage(CM) decision-support web application was originally developed by U.C. Cooperative Extension to support evapotranspiration (ET) based irrigation scheduling and nutrient management of cool-season vegetables. The model uses prescribed crop phenology curves to develop daily estimates of fractional green canopy cover (Fc) within the field. Periodic Fc observations acquired by ground-based methods or imported from NASA’s Satellite Irrigation Management Support (SIMS) can be used to adjust the prescribed timeseries for such factors as weather anomalies or non-standard agronomic practice, as needed. Fc is then converted to daily crop coefficient (fraction of reference ET) values. The crop coefficient is combined with reference evapotranspiration, collected by the California Irrigation Management Information System, to derive daily ET estimates for the given field. Irrigation runtime recommendations are issued for a given date based on total ET since the last irrigation event (less any rainfall), and corrected for distribution uniformity of the water delivery system. In this project, CM was adapted to vineyards by adding sub-models to account for early-season depletion of stored soil moisture, cover crop presence, and intentional water stress. An initial trial was performed on a Central Coast winegrape vineyard, where an eddy covariance tower measured daily ET during from May-Dec 2020. Total ET for the period showed strong agreement between the tower-based measurements (443 mm) and the CM model (431 mm). The model tended to overestimate cumulative ET during mid-June by up to 30 mm (about 15%) and later underestimated cumulative ET by as much as 65 mm (about 23%) in late September, suggesting that additional model calibration is needed to improve simulation of within-season variability. Results will be reported for trials on additional vineyard sites conducted during the 2021 season.

Evaluation↗

A parallel row-based algorithm with error control for standard-cell replacement on a hypercube multiprocessor

A new row-based parallel algorithm for standard-cell placement targeted for execution on a hypercube multiprocessor is presented. Key features of this implementation include a dynamic simulated-annealing schedule, row-partitioning of the VLSI chip image, and two novel new approaches to controlling error in parallel cell-placement algorithms; Heuristic Cell-Coloring and Adaptive (Parallel Move) Sequence Control. Heuristic Cell-Coloring identifies sets of noninteracting cells that can be moved repeatedly, and in parallel, with no buildup of error in the placement cost. Adaptive Sequence Control allows multiple parallel cell moves to take place between global cell-position updates. This feedback mechanism is based on an error bound derived analytically from the traditional annealing move-acceptance profile. Placement results are presented for real industry circuits and the performance is summarized of an implementation on the Intel iPSC/2 Hypercube. The runtime of this algorithm is 5 to 16 times faster than a previous program developed for the Hypercube, while producing equivalent quality placement. An integrated place and route program for the Intel iPSC/2 Hypercube is currently being developed.

Sargent, Jeff Scott↗

Aerial drone fleet deployment optimization with endogenous battery replacements for direct delivery of time-sensitive products

Aerial drones offer a distinct potential to reduce the delivery time and energy consumption for the delivery of time-sensitive and small products. However, there is still a need in the relevant industry to understand the performance of drone-based delivery under different business needs and drone operating conditions. We studied a drone deployment optimization problem for direct delivery of time-sensitive products with release dates to customers maintaining a specified time window. This paper presents a new mixed-integer programming model, new valid inequalities, a new greedy heuristic algorithm, and a Genetic algorithm to help business owners optimally schedule and route their drone fleet minimizing the required fleet size, the required number of additional batteries, and total energy consumption. A realistic feature of the optimization method is that instead of replacing the drone battery after each return to the depot, it keeps track of the remaining energy in the drone battery and decides on battery replacements accounting for the drone routing and the user-specified minimum required battery energy. Numerical results based on real data from drone flight tests and prepared food delivery industry provide insights into the effect of different practical drone operating parameters on the required fleet size, the required number of battery replacements, and energy consumption. Here, results demonstrate that the proposed heuristic algorithm substantially outperforms the accelerated CPLEX in runtime while sacrificing the solution quality by a small amount. Additionally, results show that using a mixed fleet of hexacopter and quadcopter drones reduces the total energy consumption by 48.52% compared to using a homogeneous fleet of only hexacopters.

Drone energy consumption↗

Discrete Event Simulation-Based Timeline Validation Using R2U2

The Gateway Vehicle Systems Manager (VSM), the top-level software control system in a distributed, hierarchical Autonomous System Management Architecture is, like most modern spacecraft software control systems, heavily data-driven. For example, schedules (timelines) will be developed on the ground and, due to the high degree of autonomy, contain complex procedures involving conditional branching, variable timing, and resource contention resolution. In order to verify that an uploaded timeline will function correctly, it is necessary to explore the feasible set of possible executions. While it is possible to test a timeline using a mission simulation, the complexity of the system and duration of a timeline limits the number of trials and therefore the test coverage. To address this problem, the VSM team is using a discrete event system model that can rapidly generate from a timeline sets of event sequences using Monte Carlo techniques. To achieve rapid and trustworthy checking of the event sequences, we use an offline version of the runtime model checking tool R2U2. This presentation describes the approach the VSM team is using to implement the discrete event simulation and evaluate event sequences using R2U2. The presentation will discuss: 1. Description of the timelines by VSM in the context of VSM operations 2. Expansion of a timeline into a sequence of atomic events 3. Adjustment, in the Monte Carlo environment, of an event sequence to account for uncertainty, external events, and failures 4. Definition of R2U2 input and mission-time linear temporal logic files 5. Generation and use of R2U2 verdict sequences 6. Lessons learned and future work

Verification↗

ATD-2 Benefits Mechanism

NASA has been developing and demonstrating a suite of decision support capabilities for integrated arrival, departure, and surface (IADS) operations in a metroplex environment. The effort is being made in three phases, under NASA’s Airspace Technology Demonstration 2 (ATD-2) sub-project, through a close partnership with the Federal Aviation Administration (FAA), air carriers, airport, and general aviation community. The Phase 1 Baseline IADS capabilities provide enhanced operational efficiency and predictability of flight operations through data exchange and integration, tactical surface metering, and automated coordination of release time of controlled flights for overhead stream insertion. The Phase 2 Fused IADS capabilities include the fusion of strategic and tactical surface metering, Atlanta Center airspace tactical scheduling, Electronic Flight Data (EFD) integration, Terminal Flight Data Manager (TFDM) Terminal Publication (TTP) prototype, and Mobile App for General Aviation (GA) community. In the Phase 2 field evaluation, strategic surface metering provides advance notice of metering and additional stability to the assigned gate holds. The users of the IADS system in Phases 1 and 2 include the personnel at Charlotte Douglas International Airport (CLT) air traffic control tower, American Airlines ramp tower, CLT terminal radar approach control (TRACON), and Washington and Atlanta Center. This document describes the ATD-2 benefits mechanism used to assess the Phases 1 and 2 IADS capabilities and field evaluation conducted at CLT since September 2017. The ATD-2 benefits mechanism mainly consists of surface metering and overhead stream insertion. This document provides detailed calculation methods of major benefit metrics, such as fuel savings, gas emissions savings, and engine runtime reduction, which can be obtained through surface metering, gate hold of Approval Request (APREQ) flights prior to pushback, and the renegotiation of release time while taxiing. As of March 31, 2020, it is estimated that 5,075,981 pounds of fuel savings and 15,634,022 pounds of CO2 emission reduction have been achieved so far, with a reduction of 3,832 hours in total engine runtime. The amount of CO2 savings is estimated to be equivalent to planting 116,254 urban trees. The pre- and post-metering comparison results using FAA’s Aviation System Performance Metrics (ASPM) data have also shown that the surface metering had no negative impact on the on-time arrival performance of both outbound and inbound flights at CLT.

ATD-2↗

ATD-2 Benefits Mechanism

NASA has been developing and demonstrating a suite of decision support capabilities for integrated arrival, departure, and surface (IADS) operations in a metroplex environment. The effort is being made in three phases, under NASA’s Airspace Technology Demonstration 2 (ATD-2) sub-project, through a close partnership with the Federal Aviation Administration (FAA), air carriers, airport, and general aviation community. The Phase 1 Baseline IADS capabilities provide enhanced operational efficiency and predictability of flight operations through data exchange and integration, tactical surface metering, and automated coordination of release time of controlled flights for overhead stream insertion. The Phase 2 Fused IADS capabilities include the fusion of strategic and tactical surface metering, Atlanta Center airspace tactical scheduling, Electronic Flight Data (EFD) integration, Terminal Flight Data Manager (TFDM) Terminal Publication (TTP) prototype, and Mobile App for General Aviation (GA) community. In the Phase 2 field evaluation, strategic surface metering provides advance notice of metering and additional stability to the assigned gate holds. The users of the IADS system in Phases 1 and 2 include the personnel at Charlotte Douglas International Airport (CLT) air traffic control tower, American Airlines ramp tower, CLT terminal radar approach control (TRACON), and Washington and Atlanta Center. This document describes the ATD-2 benefits mechanism used to assess the Phases 1 and 2 IADS capabilities and field evaluation conducted at CLT since September 2017. The ATD-2 benefits mechanism mainly consists of surface metering and overhead stream insertion. This document provides detailed calculation methods of major benefit metrics, such as fuel savings, gas emissions savings, and engine runtime reduction, which can be obtained through surface metering, gate hold of Approval Request (APREQ) flights prior to pushback, and the renegotiation of release time while taxiing. As of April 30, 2020, it is estimated that 5,097,173 pounds of fuel savings and 15,699,292 pounds of CO2 emission reduction have been achieved so far, with a reduction of 3,831 hours in total engine runtime. The amount of CO2 savings is estimated to be equivalent to planting 116,739 urban trees. The pre- and post-metering comparison results using FAA’s Aviation System Performance Metrics (ASPM) data have also shown that the surface metering had no negative impact on the on-time arrival performance of both outbound and inbound flights at CLT.

ATD-2↗