CRADL: Proxy Application for Concurrent Relaxation through Accelerated Deep Learning
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
We report strengthening in complex multicomponent systems such as solid solution alloys is controlled primarily by the dynamic interactions between dislocation lines and heterogeneously distributed solute species. Modeling of extended defect length scales in such multicomponent systems becomes prohibitively expensive, motivating the development of reduced order approaches. This work explores the application of the Concurrent Atomistic-Continuum (CAC) method to model dislocation mobility in random alloys at extended length scales. By employing recently developed average-atom interatomic potentials, the average “bulk” material response in coarse-grained regions interacts with true random solute species in the atomistic-scale domain. We demonstrate that spurious stresses in domain resolution transition regions are eliminated entirely due to the CAC formulation. Simultaneously, the key details of local stress fluctuation due to randomness in the dislocation core region are captured, and fluctuating stress smoothly decays to the long-range dislocation stress field response. Dislocation mobility calculations, for line lengths over 400 nm, are computed as a function of alloy composition in the model FeNiCr system and compared to full molecular dynamics (MD). The results capture the composition-dependent trends, while reducing degrees of freedom by nearly 40%. This approach can be readily extended to any system described by an EAM potential and facilitates the study of large-scale defect dynamics in complex solute environments to support computational alloy design.
Residual neural networks (ResNets) are a promising class of deep neural networks that have shown excellent performance for a number of learning tasks, e.g., image classification and recognition. Mathematically, ResNet architectures can be interpreted as forward Euler discretizations of a nonlinear initial value problem whose time-dependent control variables represent the weights of the neural network. Hence, training a ResNet can be cast as an optimal control problem of the associated dynamical system. For similar time-dependent optimal control problems arising in engineering applications, parallel-in-time methods have shown notable improvements in scalability. This paper demonstrates the use of those techniques for efficient and effective training of ResNets. The proposed algorithms replace the classical (sequential) forward and backward propagation through the network layers with a parallel nonlinear multigrid iteration applied to the layer domain. This adds a new dimension of parallelism across layers that is attractive when training very deep networks. From this basic idea, we derive multiple layer-parallel methods. The most efficient version employs a simultaneous optimization approach where updates to the network parameters are based on inexact gradient information in order to speed up the training process. Finally, using numerical examples from supervised classification, we demonstrate that the new approach achieves a training performance similar to that of traditional methods, but enables layer-parallelism and thus provides speedup over layer-serial methods through greater concurrency.
Previous research on adaptation to visual-motor rearrangement suggests that the central nervous system represents accurately only 1 visual-motor mapping at a time. This idea was examined in 3 experiments where subjects tracked a moving target under repeated alternations between 2 initially interfering mappings (the 'normal' mapping characteristic of computer input devices and a 108' rotation of the normal mapping). Alternation between the 2 mappings led to significant reduction in error under the rotated mapping and significant reduction in the adaptation aftereffect ordinarily caused by switching between mappings. Color as a discriminative cue, interference versus decay in adaptation aftereffect, and intermanual transfer were also examined. The results reveal a capacity for multiple concurrent visual-motor mappings, possibly controlled by a parametric process near the motor output stage of processing.
Until recently, NASA did not consider allowing computers total control of flight systems. Human operators, via hardware, have constituted the ultimate safety control. In an attempt to reduce costs, NASA has come to rely more and more heavily on computers and software to control space missions. (For example. software is now planned to control most of the operational functions of the International Space Station.) Thus the need for systematic software safety programs has become crucial for mission success. Concurrent engineering principles dictate that safety should be designed into software up front, not tested into the software after the fact. 'Cost of Quality' studies have statistics and metrics to prove the value of building quality and safety into the development cycle. Unfortunately, most software engineers are not familiar with designing for safety, and most safety engineers are not software experts. Software written to specifications which have not been safety analyzed is a major source of computer related accidents. Safer software is achieved step by step throughout the system and software life cycle. It is a process that includes requirements definition, hazard analyses, formal software inspections, safety analyses, testing, and maintenance. The greatest emphasis is placed on clearly and completely defining system and software requirements, including safety and reliability requirements. Unfortunately, development and review of requirements are the weakest link in the process. While some of the more academic methods, e.g. mathematical models, may help bring about safer software, this paper proposes the use of currently approved software methodologies, and sound software and assurance practices to show how, to a large degree, safety can be designed into software from the start. NASA's approach today is to first conduct a preliminary system hazard analysis (PHA) during the concept and planning phase of a project. This determines the overall hazard potential of the system to be built. Shortly thereafter, as the system requirements are being defined, the second iteration of hazard analyses takes place, the systems hazard analysis (SHA). During the systems requirements phase, decisions are made as to what functions of the system will be the responsibility of software. This is the most critical time to affect the safety of the software. From this point, software safety analyses as well as software engineering practices are the main focus for assuring safe software. While many of the steps proposed in this paper seem like just sound engineering practices, they are the best technical and most cost effective means to assure safe software within a safe system.
While probabilistic risk assessment (PRA) of nuclear facilities is expected to include internal and external hazards for a risk-informed and performance-based design, the current state of practice treats each hazard independently. However, such an independent treatment of hazards may not account for the correlations between different hazards and their response of and damage to the structures, systems, and components (SSCs) in a plant resulting in underestimating the overall risk. This project proposes to advance the multi-hazard PRA of nuclear facilities to more adequately evaluate concurrent hazards and contribute to an increased safety of nuclear plants. A framework for multi-hazard PRA will be developed by identifying concurrent hazard events (both internal and external) and event sequences that include interdependencies through the response of SSCs. An example application of the multi-hazard PRA framework will be demonstrated by considering a generic pressurized water reactor (PWR) subjected to seismic and internal flooding hazards. Computational models for the response of components will be developed to generated multi-hazard fragility surfaces under seismic and flooding loads. A PRA model consisting of event and fault trees will also be developed to quantify the multi-hazard risk profile and compare it with the independent hazard risk profile. Overall, by advancing the multi-hazard PRA of nuclear facilities, this project enhances nuclear safety and reduces costs by mitigating unforeseen consequences caused by correlations between concurrent hazards.
The oscillating quartz crystal viscometer has been used to investigate possible viscoelastic behavior in synthetic lubricating fluids and to obtain viscosity-pressure-temperature data for these fluids at temperatures to 300 F and pressures to 40,000 psig. The effect of pressure and temperature on the density of the test fluids was measured concurrently with the viscosity measurements. Viscoelastic behavior of one fluid, di-(2-ethylhexyl) sebacate, was observed over a range of pressures. These data were used to compute the reduced shear elastic (storage) modulus and reduced loss modulus for this fluid at atmospheric pressure and 100 F as functions of reduced frequency.
Special radiosonde soundings at 75 km spacings and 3 hour intervals provided an opportunity to learn more about mesoscale data and storm-environment interactions. Relatively small areas of intense convection produce major changes in surrounding fields of thermodynamic, kinematic, and energy variables. The Red River Valley tornado outbreak was studied. Satellite imagery and surface data were used to specify cloud information needed in the radiative heating/cooling calculations. A feasibility study for computing boundary layer winds from satellite-derived thermal data was completed. Winds obtained from TIROS-N retrievals compared very favorably with corresponding values from concurrent rawisonde thermal data, and both sets of thermally-derived winds showed good agreements with observed values.
Monthly 2.5-deg gridpoint anomalies in the Tiros-N satellite series Microwave Sounding Unit channel 2 brightness temperatures during 1979-1988 are evaluated with multiple satellites and radiosonde data for their climate temperature monitoring capability. The MSU anomalies are computed about a 10-yr mean annual cycle at each gridpoint, with the MSUs intercalibrated to a common arbitrary level. The monthly gridpoint anomaly agreement between concurrently operating satellites reveals single-satellite precision generally better than 0.07 C in the tropics and better than 0.15 C at higher latitudes. The removal from channel 2 of the temperature influence above the 30-kPa level is addressed, providing a sharper and thus potentially more useful weighting function for monitoring lower tropospheric temperatures.
There are now over one million UNIX sites and the pace at which new installations are added is steadily increasing. Along with this increase, comes a need to develop simple efficient, effective and adaptable ways of simultaneously collecting real-time diagnostic and performance data. This need exists because distributed systems can give rise to complex failure situations that are often un-identifiable with single-machine diagnostic software. The simultaneous collection of error and performance data is also important for research in failure prediction and error/performance studies. This paper introduces a portable method to concurrently collect real-time diagnostic and performance data on a distributed UNIX system. The combined diagnostic/performance data collection is implemented on a distributed multi-computer system using SUN4's as servers. The approach uses existing UNIX system facilities to gather system dependability information such as error and crash reports. In addition, performance data such as CPU utilization, disk usage, I/O transfer rate and network contention is also collected. In the future, the collected data will be used to identify dependability bottlenecks and to analyze the impact of failures on system performance.
Compared to the customary column-oriented approaches, block-oriented, distributed-memory sparse Cholesky factorization benefits from an asymptotic reduction in interprocessor communication volume and an asymptotic increase in the amount of concurrency that is exposed in the problem. Unfortunately, block-oriented approaches (specifically, the block fan-out method) have suffered from poor balance of the computational load. As a result, achieved performance can be quite low. This paper investigates the reasons for this load imbalance and proposes simple block mapping heuristics that dramatically improve it. The result is a roughly 20% increase in realized parallel factorization performance, as demonstrated by performance results from an Intel Paragon system. We have achieved performance of nearly 3.2 billion floating point operations per second with this technique on a 196-node Paragon system.
A computer program implements reference counting pointers (RCPs) that are lock-free, thread-safe, async-safe, and operational on a multiprocessor computer. RCPs are powerful and convenient means of managing heap memory in C++ software. Most prior RCP programs use locks to ensure thread safety and manage concurrency. The present program was developed in a continuing effort to explore ways of using the C++ programming language to develop safety-critical and mission- critical software. This effort includes exploration of lock-free algorithms because they offer potential to avoid some costly and difficult verification problems. Unlike previously published RCP software, the present program does not use locks (meaning that no thread can block progress on another thread): Instead, this program implements algorithms that exploit capabilities of central-processing- unit hardware so as to avoid locks. Once locks are eliminated, it becomes possible to realize the other attributes mentioned in the first sentence. In addition to the abovementioned attributes, this program offers several advantages over other RCP programs that use locks: It is smaller (and, hence, is faster and uses less memory), it is immune to priority inversion, and there is no way for it to cause a C++ exception.
High-performance computing relies on ever finer threading. Advances in processor technology include ever greater numbers of cores, hyperthreading, accelerators with integrated blocks of cores, and special vectorized instructions, all of which require more software parallelism to achieve peak performance. Traditional visualization solutions cannot support this extreme level of concurrency. Extreme scale systems require a new programming model and a fundamental change in how we design algorithms. To address these issues we created VTK-m: the visualization toolkit for multi-/many-core architectures.
Some issues in designing computers for artificial intelligence (AI) processing are discussed. These issues are divided into three levels: the representation level, the control level, and the processor level. The representation level deals with the knowledge and methods used to solve the problem and the means to represent it. The control level is concerned with the detection of dependencies and parallelism in the algorithmic and program representations of the problem, and with the synchronization and sheduling of concurrent tasks. The processor level addresses the hardware and architectural components needed to evaluate the algorithmic and program representations. Solutions for the problems of each level are illustrated by a number of representative systems. Design decisions in existing projects on AI computers are classed into top-down, bottom-up, and middle-out approaches.
Execution of heterogeneous workflows on high-performance computing (HPC) platforms present unprecedented resource management and execution coordination challenges for runtime systems. Task heterogeneity increases the complexity of resource and execution management, limiting the scalability and efficiency of workflow execution. Re-source partitioning and distribution of tasks execution over portioned re-sources promises to address those problems but we lack an experimental evaluation of its performance at scale. Here this paper provides a performance evaluation of the Process Management Interface for Exascale (PMIx) and its reference implementation PRRTE on the leadership-class HPC plat-form Summit, when integrated into a pilot-based runtime system called RADICAL-Pilot. We partition resources across multiple PRRTE Distributed Virtual Machine (DVM) environments, responsible for launching tasks via the PMIx interface. We experimentally measure the work-load execution performance in terms of task scheduling/launching rate and distribution of DVM task placement times, DVM startup and termination overheads on the Summit leadership-class HPC platform. Integrated solution with PMIx/PRRTE enables using an abstracted, standardized set of interfaces for orchestrating the launch process, dynamic process management and monitoring capabilities. It extends scaling capabilities allowing to overcome a limitation of other launching mechanisms (e.g., JSM/LSF). Explored different DVM setup configurations provide insights on DVM performance and a layout to leverage it. Our experimental results show that heterogeneous workload of 65,500 tasks on 2048 nodes, and partitioned across 32 DVMs, runs steady with resource utilization not lower than 52%. While having less concurrently executed tasks resource utilization is able to reach up to 85%, based on results of heterogeneous workload of 8200 tasks on 256 nodes and 2 DVMs.
Here, we address the calibration of a computationally expensive nuclear physics model for which derivative information with respect to the fit parameters is not readily available. Of particular interest is the performance of optimization-based training algorithms when dozens, rather than millions or more, of training data are available and when the expense of the model places limitations on the number of concurrent model evaluations that can be performed. As a case study, we consider the Fayans energy density functional model, which has characteristics similar to many model fitting and calibration problems in nuclear physics. We analyze hyperparameter tuning considerations and variability associated with stochastic optimization algorithms and illustrate considerations for tuning in different computational settings.
Many elementary deformation processes in metals involve the motion of dislocations. The planes of glide and specific processes dislocations prefer depend heavily on their atomic core structures. Atomistic simulations are desirable for dislocation modeling but their application to even sub-micron scale problems is in general computationally costly. Accordingly, continuum-based approaches, such as the phase-field microelasticity, phase-field dislocation dynamics (PFDD), generalized Peierls–Nabarro (GPN) models, and the concurrent atomistic–continuum (CAC) method, have attracted increasing attention in the field of dislocation modeling because they well represent both short-range cores interactions and long-range stress fields of dislocations. To better understand their similarities and differences, it is useful to compare these methods in the context of benchmark simulations and predictions. In this paper, we apply the CAC method and different PFDD variants – one of them is equivalent to a GPN model – to simulate an extended (i.e., dissociated) dislocation in Al with initially pure edge or pure screw character in terms of the disregistry. CAC and discrete forms of PFDD are also employed to calculate the Peierls stress. By conducting comprehensive convergence studies, we quantify the dependence of these measures on time/grid resolution and simulation cell size. Several important but often overlooked differences between PFDD/GPN variants are clarified. In conclusion, our work sheds light on the advantages and limitations of each method, as well as the path towards enabling them to effectively model complex dislocation processes at larger length scales.
Conversion plateaus rapidly in radical photopolymerizations (RPPs) following discontinuation of irradiation due to rapid termination of reactive radicals, which restricts the wider use of RPPs in applications that involve nonuniform light access including those with attenuated light transmission or irregular surfaces. Based on our recent report of a radical dark-curing photoinitiator (DCPI) that continues polymerization beyond the cessation of irradiation by enabling latent redox initiation with photo-released amine in the presence of a suitable oxidant, we developed a new DCPI with an absorption spectrum that extends well into the visible range. Our design process involved a series of computational investigations of candidate molecules, including a systematic study of substituents and their position-dependent effects on absorption characteristics, electronic transitions, and the photochemical mechanism and its associated energetics. Our quantum chemical computations identified the target compound 5,7-dimethoxy-6-bromo-3-aroylcoumarin-DMPT/BPh4 and predicted that it would facilitate the dark-curing mechanism by concurrent photo-radical generation and photo-induced release of an efficient redox reductant under visible irradiation. This reductant-tethered chromophore was then synthesized and optically characterized with UV–vis spectroscopy that revealed its strong visible-light absorption with a molar absorptivity of 5710 M–1 cm–1 at 405 nm and 50 M–1 cm–1 at 455 nm. We then demonstrated extensive dark-curing of >35% additional conversion over 25 min following brief activation of the shelf-stable one-part system by irradiation with a 455 nm LED that was ceased at 20% conversion. In contrast, shuttering irradiation of the control formulation at that same point resulted in immediate cessation of conversion, which plateaued at 20%. We determined a remarkable initiator efficiency of 2.82 that results from the additional redox-generated radicals with a 77% photo-reductant generation quantum yield. The combination of superior photo- and dark-curing efficiencies of this new visible DCPI is expected to open new application opportunities in RPP, especially those involving resins that are highly light attenuating, surfaces that possess irregular features that produce uneven irradiance, and production lines where continued dark-curing downstream of the light activation step enhances line efficiencies.