Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Application Development”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Roadmap for Photonics with 2D Materials

Triggered by advances in atomic-layer exfoliation and growth techniques, along with the identification of a wide range of extraordinary physical properties in self-standing films consisting of one or a few atomic layers, two-dimensional (2D) materials such as graphene, transition metal dichalcogenides (TMDs), and other van der Waals (vdW) crystals now constitute a broad research field expanding in multiple directions through the combination of layer stacking and twisting, nanofabrication, surface-science methods, and integration into nanostructured environments. Photonics encompasses a multidisciplinary subset of those directions, where 2D materials contribute remarkable nonlinearities, long-lived and ultraconfined polaritons, strong excitons, topological and chiral effects, susceptibility to external stimuli, accessibility, robustness, and a completely new range of photonic materials based on layer stacking, gating, and the formation of moiré patterns. These properties are being leveraged to develop applications in electro-optical modulation, light emission and detection, imaging and metasurfaces, integrated optics, sensing, and quantum physics across a broad spectral range extending from the far-infrared to the ultraviolet, as well as enabling hybridization with spin and momentum textures of electronic band structures and magnetic degrees of freedom. The rapid expansion of photonics with 2D materials as a dynamic research arena is yielding breakthroughs, which this Roadmap summarizes while identifying challenges and opportunities for future goals and how to meet them through a wide collection of topical sections prepared by leading practitioners.

2D materials↗

Crystal phases of charged interlayer excitons in van der Waals heterostructures

Throughout the years, strongly correlated coherent states of excitons have been the subject of intense theoretical and experimental studies. This topic has recently boomed due to new emerging quantum materials such as van der Waals (vdW) bound atomically thin layers of transition metal dichalcogenides (TMDs). We analyze the collective properties of charged interlayer excitons observed recently in bilayer TMD heterostructures. We predict strongly correlated phases—crystal and Wigner crystal—that can be selectively realized with TMD bilayers of properly chosen electron-hole effective masses by just varying their interlayer separation distance. Our results can be used for nonlinear coherent control, charge transport and spinoptronics application development with quantum vdW heterostuctures.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Battle of the Defaults: Extracting Performance Characteristics of HDF5 under Production Load

Popular parallel I/O libraries, such as HDF5, provide tuning parameters to obtain superior performance. However, the selection of effective parameters on production systems is complex due to the interdependence of I/O software and file system layers. Hence, application developers typically use the default parameters and often experience poor I/O performance. This work conducts a benchmarking-based analysis on the HDF5 behaviors with a wide variety of I/O patterns to extract performance characteristics under the production workload. To make the analysis well controlled, we exercise I/O benchmarks on POSIX-IO, MPI-IO, and HDF5 using the same I/O patterns and in the same jobs. To address high performance variability in production environments, we repeat the benchmarks across I/O patterns, storage devices, and time intervals. Based on the results, we identified consistent HDF5 behaviors that appropriate configurations and operations on dataset layout and file-metadata placement can improve performance significantly. We apply our findings and evaluate the tuned I/O library on two supercomputers: Summit and Cori. The results show that our tuned parameters can achieve more than 10× I/O performance speedup than that with default parameters on both systems, suggesting the effectiveness, stability, and generality of our solution.

Xie, Bing↗

IRIS-MEMFLOW: Data Flow-Enabled Portable Memory Orchestration in IRIS Runtime for Diverse Heterogeneity

Task-based programming models and execution paradigms provide a means to decompose a computation by expressing it as a graph in which each node represents a specific computation operating on memory objects and the edges define the dependencies in the execution flow. In this execution model, independent nodes in the graph can be executed concurrently in different computing devices, making it suitable for heterogeneous systems in which computing devices with different architectures coexist. However, careful memory orchestration across heterogeneous devices is needed because copies of the same memory object may reside in multiple devices during execution. Manually ensuring such an orchestration is quite challenging. Not only must an application developer guard against race conditions, but they must also optimize data movement between the host and devices because unnecessary data movement significantly impacts performance. To mitigate these challenges, we enhance the IRIS heterogeneous runtime and introduce IRIS-MEMFLOW–a data flow–enabled portable memory abstraction for seamlessly orchestrating memory in diverse heterogeneous computing environments. By using data-flow analysis, IRIS-MEMFLOW guards against race conditions while multiple heterogeneous devices access memory objects. IRIS-MEMFLOW also optimizes data movement between the host and devices without manual intervention. As a result, IRIS provides improved programming productivity, performance, and portability for multidevice heterogeneous executions in high-performance computing and cloud systems that run diverse architectures from different vendors. The efficacy of IRIS-MEMFLOW is evaluated through experiments that show its capability in terms of programming productivity, multidevice heterogeneity, portability, and low overhead versus the state of the art.

Monil, M. A. H. [ORNL] (ORCID:0000000334194037)↗

Performance Portability in the Exascale Computing Project: Exploration Through a Panel Series

Performance portability is a critical issue for the Exascale Computing Project (ECP) because of nontrivial architectural differences between machines available today and those expected at exascale. Many ECP project teams are working toward performance portability, and would expect to benefit from sharing lessons learned, identifying gaps, and discovering opportunities for partnerships. To facilitate this communication, the IDEAS-ECP project partnered with the three focus areas of ECP (application development, software technology, and hardware and integration), and Department of Energy computing facilities, to lead a series of panel discussions on performance portability. The panels were organized around broadly common themes of algorithmic and data locality challenges. In this article, we describe the panel series, its objectives, and perspectives from the various areas of the project. Finally, we also discuss use cases that are distinctive, as well as conclusions drawn from the collective experience of the participants.

97 MATHEMATICS AND COMPUTING↗

Bridging Python to Silicon: The SODA Toolchain

Systems performing scientific computing, data analysis, and machine learning tasks have a growing demand for application-specific accelerators that can provide high computational performance while meeting strict size and power requirements. However, the algorithms and applications that need to be accelerated are evolving at a rate that is incompatible with manual design processes based on hardware description languages. Agile hardware design tools based on compiler techniques can help by quickly producing an application-specific integrated circuit (ASIC) accelerator starting from a high-level algorithmic description. Here, we present the software-defined accelerator (SODA) synthesizer, a modular and open-source hardware compiler that provides automated end-to-end synthesis from high-level software frameworks to ASIC implementation, relying on multilevel representations to progressively lower and optimize the input code. Our approach does not require the application developer to write any register-transfer level code, and it is able to reach up to 364 giga floating point operations per second (GFLOPS)/W efficiency (32-bit precision) on typical convolutional neural network operators.

97 MATHEMATICS AND COMPUTING↗

A Framework for Integrating Quantum Simulation and High Performance Computing

Scientific applications are starting to explore the viability of quantum computing. This exploration typically begins with quantum simulations that can run on existing classical platforms, albeit without the performance advantages of real quantum resources. In the context of high-performance computing (HPC), the incorporation of simulation software can often take advantage of the powerful resources to help scale-up the simulation size. The configuration, installation and operation of these quantum simulation packages on HPC resources can often be rather daunting and increases friction for experimentation by scientific application developers. We describe a framework to help streamline access to quantum simulation software running on HPC resources. This includes an interface for circuit-based quantum computing tasks, as well as the necessary resource management infrastructure to make effective use of the underlying HPC resources. The primary contributions of this work include a classification of different usage models for quantum simulation in an HPC context, a review of the software architecture for our approach and a detailed description of the prototype implementation to experiment with these ideas using two different simulators (TNQVM & NWQ-Sim). We include initial experimental results running on the Frontier supercomputer at the Oak Ridge Leadership Computing Facility (OLCF) using a synthetic workload generated via the SupermarQ quantum benchmarking framework.

Shehata, Amir [ORNL] (ORCID:0000000224531426)↗

Cell‐type‐specific transcriptomics uncovers spatial regulatory networks in bioenergy sorghum stems

SUMMARY Bioenergy sorghum is a low‐input, drought‐resilient, deep‐rooting annual crop that has high biomass yield potential enabling the sustainable production of biofuels, biopower, and bioproducts. Bioenergy sorghum's 4–5 m stems account for ~80% of the harvested biomass. Stems accumulate high levels of sucrose that could be used to synthesize bioethanol and useful biopolymers if information about cell‐type gene expression and regulation in stems was available to enable engineering. To obtain this information, laser capture microdissection was used to isolate and collect transcriptome profiles from five major cell types that are present in stems of the sweet sorghum Wray. Transcriptome analysis identified genes with cell‐type‐specific and cell‐preferred expression patterns that reflect the distinct metabolic, transport, and regulatory functions of each cell type. Analysis of cell‐type‐specific gene regulatory networks (GRNs) revealed that unique transcription factor families contribute to distinct regulatory landscapes, where regulation is organized through various modes and identifiable network motifs. Cell‐specific transcriptome data was combined with known secondary cell wall (SCW) networks to identify the GRNs that differentially activate SCW formation in vascular sclerenchyma and epidermal cells. The spatial transcriptomic dataset provides a valuable source of information about the function of different sorghum cell types and GRNs that will enable the engineering of bioenergy sorghum stems, and an interactive web application developed during this project will allow easy access and exploration of the data ( https://mc‐lab.shinyapps.io/lcm‐dataset/ ).

09 BIOMASS FUELS↗

Preparation and Characterization Methods of Thin Layer Samples for Standoff Detection

Detection of analytes deposited on surfaces is crucial for many applications: Development of methods to prepare thin layers (e.g. ~5 to 100 µm) is important for both system design and field studies. In this work, solid and liquid analytes were deposited on painted and bare substrates including aluminum, glass, plastic, and concrete using an ExactaCoat ultrasonic spray coater. Laboratory hemispherical reflectance (HRF) spectra were collected for samples with different layer thicknesses so as to characterize both the composition and layer thickness. Preliminary results demonstrate that to prepare homogenous layers on surfaces, parameters such as substrate type, analyte solubility, vapor pressure, paint color, surface porosity, and surface roughness are all important. Liquid chemicals posed several issues during deposition: Diisopropyl methyl phosphonate evaporated from surfaces more quickly than the other chemicals and was thus not detected in the HRF experiments. Less volatile liquids, such as tributylphosphate, remained on the surface for the duration of the test, but a uniform layer thickness could not be obtained as the liquid pooled to one side when mounted at an angle. The deposition of solids (e.g., acetaminophen, caffeine and methylphosphonic acid) from volatile solvents such as chloroform also proved problematic due to streaking caused by rapid solvent evaporation. Solids deposited from ethanol, however, worked well on bare substrates. For most samples plotting the integrated infrared band strength vs. surface thicknesses showed a linear relationship, confirming that the surface loading can be controlled by programming the concentration and the number of passes on the ultrasonic sprayer.

Thin layer, deposition, infrared standoff, Hemisph↗

Exploring Architectural-Aware Affinity Policies in Modern HPC Runtimes

Modern commodity and High-Performance Computing (HPC) systems are evolving with complex CPU architectures. These architectures now feature higher core and NUMA domain counts and implement features such as hyperthreading. When considering significant differences in hardware configurations, library availability, and hardware-tailored system/software stacks, which could substantially vary from one system to another, performance portability is hard to achieve. Throughout the years, this trend resulted in an increasingly high burden on application developers to fine-tune their workloads for each architecture. This work explores how hardware-dependent aspects such as locality/process/thread affinity affect performance in modern CPU architectures. We focus our study on the Global Memory and Threading (GMT) distributed runtime system as a representative of Partitioned Global Address Space (PGAS) software stacks commonly adopted for productivity. In particular, to appreciate performance implications, we evaluate GMT’s thread affinity policies, and, introduce two new ones which exploit architectural awareness. Finally, we explore alternative NUMA configurations via different process bindings and perform a scalability study on three HPC clusters with varying CPU architectures and NUMA layouts. Our analysis indicates that more complex architectures are more affected by affinity and binding policies and highlights the importance of setting proper runtime configurations to achieve superior performance.

Di Dio Lavore, Ian↗

Flibcpp

Flibcpp is a library for use by application developers that provides native Fortran interfaces to existing high-quality algorithms and containers implemented in C++ and available on all modern computer systems.

Johnson, SethR [Oak Ridge National Laboratory] (00↗

Building Efficiency Targeting Tool for Energy Retrofits (BETTER) Web Application (BETTER Web App) v1.0

The Building Efficiency Targeting Tool for Energy Retrofits (BETTER) V.1.0 web application identifies cost-saving energy and emissions reductions in buildings and portfolios, without site visits or complex modeling. With minimal data entry, BETTER benchmarks a building's or portfolio's energy use against peers; quantifies energy, cost, and greenhouse gas (GHG) reduction potential; and recommends energy efficiency measures for individual buildings or portfolios, targeting specific energy savings levels. The source code of BETTER's modular, cross-platform analytical engine has previously been disclosed (2019-001) and is available on GitHub and can be adopted, redeveloped, and redistributed freely under an open-source license, allowing users to incorporate BETTER's analytical capabilities into their own software platforms and tools. The BETTER V.1.0 web application, developed with the Django web-framework and the Model-View-Controller (MVC) architecture being disclosed here, provides a graphical user interface (GUI) for any user to view the software tutorials, download and upload a data entry template, configure and run the BETTER analyses, and view the final analytical reports.

Szum, Carolyn↗

Ingest

Ingest is an application that allows project owners to solicit data from users and require them to fill out metadata along with those uploads. It is an application developed in Elixir/Phoenix and is a web based platform.

Darrington, JohnW.↗

GridOPTICS/GridPACK

GridPACK is a software framework consisting of a set of modules designed to simplify the development of programs that model the power grid and run on parallel, high performance computing platforms. It also contains several fully developed applications, including powerflow, dynamic simulation, state estimation, Kalman filter analysis (dynamic state estimation), contingency analysis and real time path rating. These applications can be used either standalone or as components in more complicated workflows that combine several different types of application together. The framework modules are available as a combination of libraries and software templates and consist of components for setting up and distributing power grid networks, support for modeling the behavior of individual buses and branches in the network, converting the network models to the corresponding algebraic equations, and parallel routines for manipulating and solving large algebraic systems. The framework also contains a module for distributing tasks evenly amongst computing resources, even if individual tasks vary widely in their execution times. Additional modules support input and output, basic statistical analysis of contingency based calculations, distributed data structures, as well as basic profiling and error management.

Palmer, Bruce↗

Fayda

The Fayda application developed as a part of the Reaction Roulette m/q LDRD project (76006) primarily exists as a data visualization, analysis and predictive platform for data collected using atomic tandem inductively coupled plasma mass spectrometry (ICP-MS/MS)

Harouaka, Khadouja↗

VerifyIO: Verifying Adherence to Parallel I/O Consistency Semantics

VerifyIO is a tool designed for verifying I/O consistency semantics in High-Performance Computing (HPC) applications. It addresses the challenges of ensuring correctness and portability across different I/O consistency models, such as POSIX, Commit, Session, and MPI-IO. By analyzing execution traces, detecting conflicts, and verifying synchronization adherence, VerifyIO provides actionable insights for both application developers and I/O library designers.

Wang, Chen [Lawrence Livermore National Laboratory↗

Coupling of regional geophysics and local soil-structure models in the EQSIM fault-to-structure earthquake simulation framework

Accurate understanding and quantification of the risk to critical infrastructure posed by future large earthquakes continues to be a very challenging problem. Earthquake phenomena are quite complex and traditional approaches to predicting ground motions for future earthquake events have historically been empirically based whereby measured ground motion data from historical earthquakes are homogenized into a common data set and the ground motions for future postulated earthquakes are probabilistically derived based on the historical observations. This procedure has recognized significant limitations, principally due to the fact that earthquake ground motions tend to be dictated by the particular earthquake fault rupture and geologic conditions at a given site and are thus very site-specific. Historical earthquakes recorded at different locations are often only marginally representative. There has been strong and increasing interest in utilizing large-scale, physics-based regional simulations to advance the ability to accurately predict ground motions and associated infrastructure response. However, the computational requirements for simulations at frequencies of engineering interest have proven a major barrier to employing regional scale simulations. In a U.S. Department of Energy Exascale Computing Initiative project, the EQSIM application development is underway to create a framework for fault-to-structure simulations. This framework is being prepared to exploit emerging exascale platforms in order to overcome computational limitations. This article presents the essential methodology and computational workflow employed in EQSIM to couple regional-scale geophysics models with local soil-structure models to achieve a fully integrated, complete fault-to-structure simulation framework. Here, the computational workflow, accuracy and performance of the coupling methodology are illustrated through example fault-to-structure simulations.

97 MATHEMATICS AND COMPUTING↗

Electrochemically modulated single-molecule localization microscopy for in vitro imaging cytoskeletal protein structures

A new concept of electrochemically modulated single-molecule localization super-resolution imaging is developed. Applications of single-molecule localization super-resolution microscopy have been limited due to insufficient availability of qualified fluorophores with favorable low duty cycles. The key for the new concept is that the “On” state of a redox-active fluorophore with unfavorable high duty cycle could be driven to “Off” state by electrochemical potential modulation and thus become available for single-molecule localization imaging. The new concept was carried out using redox-active cresyl violet with unfavorable high duty cycle as a model fluorophore by synchronizing electrochemical potential scanning with a single-molecule localization microscope. The two cytoskeletal protein structures, the microtubules from porcine brain and the actins from rabbit muscle, were selected as the model target structures for the conceptual imaging in vitro. The super-resolution images of microtubules and actins were obtained from precise single-molecule localizations determined by modulating the On/Off states of single fluorophore molecules on the cytoskeletal proteins via electrochemical potential scanning. Importantly, this method could allow more fluorophores even with unfavorable photophysical properties to become available for a wider and more extensive application of single-molecule localization microscopy.

electrochemical modulation↗