Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Spot: A Programming Language for Verified Flight Software

The C programming language is widely used for programming space flight software and other safety-critical real time systems. C, however, is far from ideal for this purpose: as is well known, it is both low-level and unsafe. This paper describes Spot, a language derived from C for programming space flight systems. Spot aims to maintain compatibility with existing C code while improving the language and supporting verification with the SPIN model checker. The major features of Spot include actor-based concurrency, distributed state with message passing and transactional updates, and annotations for testing and verification. Spot also supports domain-specific annotations for managing spacecraft state, e.g., communicating telemetry information to the ground. We describe the motivation and design rationale for Spot, give an overview of the design, provide examples of Spot's capabilities, and discuss the current status of the implementation.

validation↗

Visualizing Corridors in Terminal Airspace using Trajectory Clustering

Context: Advances in battery and automation technology have made routine air taxi and cargo transport in urban areas a business model that can be attained by emerging aviation innovators. The community vision and work to enable these novel operations is discussed using the term ‘Urban Air Mobility’ or UAM. Small, piloted, airspace vehicles that fly with a few passengers do operate in urban areas today, and these vehicles can be studied as an early proxy for this future UAM traffic. Aim: We seek to identify corridors already in daily operation and their properties. Method: We applied DBSCAN and HDBSCAN to Dallas Forth-Worth TRACON flight data to identify corridors in use, their density, and devised a method to annotate landing sites used in these corridors with site metadata. Results: While DBSCAN was unable to group similar trajectories, we we were able to successfully identify corridors using HDBSCAN, measure their density and annotate them. Conclusion: The applied method can successfully identify corridors in daily operation with additional metadata to help domain expert understand the intent of UAM corridors.

UAM, Trajectory, TRACON, Clustering, DBSCAN, HDBSC↗

Visualizing Corridors in Terminal Airspace Using Trajectory Clustering

Context: Advances in battery and automation technology have made routine air taxi and cargo transport in urban areas a business model that can be attained by emerging aviation innovators. The community vision and work to enable these novel operations is discussed using the term ‘Urban Air Mobility’ or UAM. Small, piloted, airspace vehicles that fly with a few passengers do operate in urban areas today, and these vehicles can be studied as an early proxy for this future UAM traffic. Aim: We seek to identify corridors already in daily operation and their properties. Method: We applied DBSCAN and HDBSCAN to Dallas Forth-Worth TRACON flight data to identify corridors in use, their density, and devised a method to annotate landing sites used in these corridors with site metadata. Results: While DBSCAN was unable to group similar trajectories, we we were able to successfully identify corridors using HDBSCAN, measure their density and annotate them. Conclusion: The applied method can successfully identify corridors in daily operation with additional metadata to help domain expert understand the intent of UAM corridors.

UAM Trajectory, TRACON, Clustering, DBSCAN, HDBSCA↗

Lunar Instrument Data Integration Into the Virtual Reality Mission Simulation System for Decision Communication and Situational Awareness

In situ resource utilization (ISRU) technologies are a key advancement required to make human habitation on the Moon and Mars viable. The upcoming Volatiles Investigating Polar Exploration Rover (VIPER) mission will provide crucial correlations between volatiles and lunar geology to understand the water content available for ISRU on the moon. The mission will require the coordination of multi-disciplinary teams across the country making real time decisions based on rover instrument data. The virtual reality Mission Simulation System (vMSS) is a virtual reality platform designed at MIT by the Resource Exploration and Science of our Cosmic Environment (RESOURCE) team to provide teams with a collaboration interface for planetary missions like VIPER. Herein we determine the integration pathway for analog based datasets that are examples of VIPER's two main instruments, the near-infrared volatile spectrometer subsystem (NIRVSS) and the neutron spectrometer subsystem (NSS), into vMSS to provide the most valuable visualization tools. Focusing on improving situational awareness, decision making, reducing task load and incorporating comments from scientists working previous analogs and on the current VIPER mission, we recommend critical elements to implement into vMSS and the best approaches for data visualization. We present a review of relevant analogs and state of the art mission software. We have developed a design concept and path to flight of analysed instrument data integrated with data maps that allow for virtual manipulation and annotation between non-co-located team members. We focus on pre-mission mapping of a priori data for improved situational awareness, layering of analysed instrument data, correlative mapping and interactive capabilities for in-mission decision making, as well as archiving and annotation tools for post-mission analysis. Finally, we lay out the roadmap for the future development of immersive sample site visualization capabilities and the use of integrated instrument data in vMSS with automated temporal and geospatial planning.

Cody Alison Paige↗

A Natural Language Understanding Approach for Digitizing Aircraft Ground Taxi Instructions

Recent Advancements in spoken language processing technologies have enabled the development of reliable automation tools to assist air traffic control (ATC) operations. These advancements present a unique opportunity to strategically implement digital taxi instructions for aircraft movement on the ground. Digital taxi instructions issue taxiing procedures to pilots as textual or graphical instructions. There are several benefits of digital instructions, including a reduction in radio congestion, elimination of communication errors, and improved aircraft monitoring capabilities. Natural language understanding (NLU) models can extract digital taxi instructions from verbal instructions issued by air traffic controllers. This capability enables the implementation of a digital taxi communication framework with minimal changes to the existing air traffic controller operational environment. We explore a novel application for NLU: automatically generating digital taxi instructions from air traffic controller speech. We describe the development of an annotation scheme to represent (aircraft) ground traffic communications in the US National Airspace System (NAS).Our annotation scheme uses intent classification and slot filling to extract taxi instructions, enabling NLU models to leverage syntactic information. Several neural network models were trained to categorize the controller’s intent and label information relevant to his or her intent.Our research demonstrates that it is feasible to use NLU to automatically generate digitaltaxi instructions, suggesting that it is a powerful tool for the implementation of digital taxi communications.

Hillel Steinmetz↗

A Natural Language Understanding Approach for Digitizing Aircraft Ground Taxi Instructions

Advancements in natural language processing (NLP) technologies offer a unique opportunity to furnish aircraft crews, primarily pilots, with digital instructions for taxiing operations. Digital taxi instructions, delivered either as text or graphics, can streamline taxiing procedures, thereby reducing radio congestion, minimizing communication errors, and enhancing aircraft monitoring. Techniques used for natural language understanding (NLU), a subset of NLP focused on machine comprehension of natural language, can extract taxi instructions directly from verbal radio communications. This capability paves the way for implementing a digital taxi communication framework with minimal adjustments to the existing air traffic controller operations. This paper delves into a novel application of NLU: the automated generation of digital taxi instructions from air traffic controller speech. We detail the development of an annotation scheme to represent aircraft ground traffic communications within the US National Airspace System (NAS), employing intent classification (IC) and slot filling (SF) to extract taxi instructions using NLU models. Several neural network models were trained on a dataset annotated with our scheme, achieving notable accuracy and F1 scores. Our research demonstrates the feasibility of using NLU to automatically generate digital taxi instructions, showcasing its potential to streamline the implementation of digital taxi communications.

LSTM↗

A Natural Language Understanding Approach for Digitizing Aircraft Ground Taxi Instructions

Advancements in natural language processing (NLP) technologies offer a unique opportunity to furnish aircraft crews, primarily pilots, with digital instructions for taxiing operations. Digital taxi instructions, delivered either as text or graphics, can streamline taxiing procedures, thereby reducing radio congestion, minimizing communication errors, and enhancing aircraft monitoring. Techniques used for natural language understanding (NLU), a subset of NLP focused on machine comprehension of natural language, can extract taxi instructions directly from verbal radio communications. This capability paves the way for implementing a digital taxi communication framework with minimal adjustments to the existing air traffic controller operations. This paper delves into a novel application of NLU: the automated generation of digital taxi instructions from air traffic controller speech. We detail the development of an annotation scheme to represent aircraft ground traffic communications within the US National Airspace System (NAS), employing intent classification (IC) and slot filling (SF) to extract taxi instructions using NLU models. Several neural network models were trained on a dataset annotated with our scheme, achieving notable accuracy and 𝐹1 scores. Our research demonstrates the feasibility of using NLU to automatically generate digital taxi instructions, showcasing its potential to streamline the implementation of digital taxi communications.

ATC↗

Cleanroom Microbes Survive Drying, Vacuum, and Proton Irradiation

Introduction : The goal of planetary protection at NASA is to mitigate the risk of contaminating sensitive target bodies with biological life. While many cleaning procedures have been put in place to reduce bioburden on spacecraft, microbes are experts at evolving to survive harsh conditions. Specifically, the dry, low-nutrient environment of a cleanroom (commonly used for assembly of spacecraft) can represent an environment where extremophiles can survive. Methods : Scientists at NASA MSFC wished to gather a snapshot of the microbial population within a variety of cleanrooms on site. A study was undertaken to collect air, surface, and floor samples from clean-rooms and isolate unique morphologies. From this study, 95 isolates were collected and saved in a microbial library. About 86% of these were identified at least to a genus level. Following identification, 24 microbes were selected, based on a literature review, as potential extremophiles. These were grown in liquid cultures, diluted to a set optical density, washed with water, and then applied to a sterilized Kapton coupon. Droplets were allowed to dry overnight in a biosafety cabinet. Coupons were then installed in a pelletron and pumped down to high vacuum (~1E-6 Torr). Samples were then subjected 100 keV protons at a fluence of 2x10 15 p+/cm 2 up to 4x10 15 p+/cm 2 . Following exposure, samples were returned to the microbiology lab where they were pro-cessed by submerging in water, vortexing, and then plating either droplets or spread plates. Recovery data collected was qualitative with a ranking or +, minor, or – for growth. Some selected radiotolerant strains were sequenced using the Illumina sequencing platform. The resulting genomes were annotated with the Rapid Annotations using Subsystems Technology (RAST) server and analyzed for conserved and unique stress response relevant genomic signatures to identify clues related to specific tolerances. Results and Discussion : After five rounds of proton radiation, we narrowed our isolates to five, non-spore forming bacteria that demonstrated survival: Arthrobacter koreensis, Paenarthrobacter nitroguajacolicus, Mycetocola manganoxydans , and an Erwinia sp. Furthermore, we exposed these four microbes to 254 nm wavelength light at an intensity of 80 W/m 2 at a distance of ~18 cm for 10 minutes. Only A. koreensis demonstrated survival following UV exposure. Finally, we performed whole genome sequencing on the four strains to look for genetic markers of stress resistance. When we compared the genomes of the four strains, we found that genes coding for GGDEF and EAL domains with PAS/PAC sensors were only found in A. koreensis . These domains, modulated by PAS/PAC sensors, are hypothesized to facilitate survival under drying, desiccation, and proton irradiation. Drying and Desiccation : PAS domains sense hydration changes and modulate GGDEF and EAL domain activity to adjust c-di-GMP levels, enhancing resistance to desiccation. For instance, in Pseudomonas aeruginosa , the PAS domain of RbdA modulates activity under varying hydration conditions, affecting stress responses [1]. Proton Irradiation : Proton irradiation causes oxidative stress, leading to ROS generation. PAS domains detect this stress and modulate GGDEF and EAL domains to manage oxidative stress responses. In Shewanella , EAL domain proteins modulated by PAS sensors help bacteria adapt to extreme conditions [2]. These genes upregulate other stress response genes, protecting membrane function, protein stability, DNA repair, and antioxidant defenses. The modulation of c-di-GMP by PAS domains is crucial for bacterial adaptation to stress conditions, enabling dynamic physio-logical adjustments [3]. Understanding these mechanisms provides insights into bacterial stress responses and strategies for controlling bacterial growth [4]. Conclusions : These findings indicate that clean-rooms harbor extremophile microbes that may be able to survive conditions in deep space. Furthermore, while we identified certain stress-response genes that may be at least partly responsible for the phenotypes observed in this study, there are likely unidentified genes or characteristics about A. koreensis , and other bacteria, that may allow them to survive in harsh environments. Future studies will focus on identifying these unknown genes and characteristics, further elucidating the mechanisms of extremophile survival and potentially informing the development of new biotechnologies for space exploration and other extreme environments.

Chelsi Cassilly↗

Contextualizing Air Traffic Management Conversations using Natural Language Understanding

Efficient management of air traffic and mitigation of delays depend on extracting actionable information from unstructured data, such as dialogues from the Federal Aviation Administration’s (FAA’s) Air Traffic Control System Command Center (ATCSCC) telecons. This study presents a pipeline utilizing Natural Language Processing (NLP) methods for Intent Classification (IC) and Slot Filling (SF) to identify and extract Traffic Management Initiatives (TMIs) from aviation-specific dialogues. We leveraged DeBERTa, a pre-trained transformer model, and fine-tuned it to the nuances of the aviation domain. Despite challenges posed by annotation complexities, the IC model achieved promising results with a weighted average F1-score of 0.81. Our results are close to those of human annotators, which demonstrates the model’s strong alignment with human-level performance. The SF model also showed strong performance, achieving a weighted F1-score of 0.97, which demonstrates its effectiveness in accurately predicting key slots. Our analysis revealed limitations in handling less frequent intents and slot labels due to data sparsity, motivating future efforts to adopt joint IC-SF modeling and data augmentation strategies. This research highlights the potential of domain-specific NLP to streamline decision-making in the aviation industry and improve the management of TMIs.

Air Traffic Control Management↗

Interactive Visualizations for Communicating Compliance and Sustainability Statements in Annual Site Environmental Reports - 20419

The annual reports of organizations are a recurring type of public disclosure that communicates what an organization has done and/or plans to do. Worldwide, organizations use annual reports as powerful instruments for communication that is rich in narrative content and visuals. From the perspective of the leadership of an organization, statutory annual reports are communication devices to stakeholders. By analogy, from the perspective of stakeholders, the annual reports are learning devices that presents stakeholders with contents to generate knowledge about an organization. We have chosen the online accessible U.S. Department of Energy (DOE) Annual Site Environmental Reports (ASERs) as a collection of resources for communicating and learning the environmental management strategies on environmental compliance and environmental sustainability. Environmental Compliance 'consists of regulatory compliance and monitoring programs that implement federal, state, and local requirements, agreements, and permits.' Environmental Sustainability 'promotes and integrates initiatives such as energy and natural resource conservation, waste minimization, green remediation, and the use of sustainable products and services.' Interactive visualizations of textual data designed with visual analytics software can allow stakeholders to interact with statements on the environmental management strategies by performing diverse actions that promote learning. These actions include annotating, comparing, filtering, navigating, selecting and sharing relevant statements. We expect that the possibility of these multiple actions in an interactive manner help stakeholders to learn more effectively about the environmental strategies at U.S. DOE sites. The current case study for the project is a collection of 1,894 annotated (value-added) statements (sentence text) from the eight chapters of the 2017 Savanah River Site (SRS) annual site environmental report. The preliminary interactive visualizations are available online. (authors)

54 ENVIRONMENTAL SCIENCES↗

GOLEM: GOld standard for Learning and Evaluation of Motifs

Motifs are distinctive, recurring, widely used idiom-like words or phrases, often originating from folklore, whose meaning is anchored in a narrative and have a significance as communicative devices across a wide range of media, including news, literature, and propaganda. Many motifs concisely imply a large constellation of culturally relevant information, and their broad usage suggests their cognitive importance as touchstones of cultural knowledge. As such, their detection is a step towards culturally aware natural language processing. We present GOLEM (GOld standard for Learning and Evaluation of Motifs) a dataset of English news articles, opinion pieces, and broadcast transcripts annotated for motific information. The dataset identifies 25,737 motif candidates across 34 motif types drawn from three cultural or national groups: Jewish, Irish, and Puerto Rican. The dataset contains 2,024,141 words split into 25,737 text snippets drawn from 8,073 articles. Each motif candidate is labeled according to a scheme which identifies the type of usage (motific, referential, eponymic, or unrelated), resulting in 1,743 actual motific instances in the data. Annotation was performed by individuals identifying as members of each group and achieved a Fleiss’ kappa (?) of > 0.55. In addition to the data, we demonstrate that classification of the candidate type is a challenging task for Large Language Models (LLMs) using a few-shot approach; recent models such as T5, FLAN-T5, GPT-2, and Llama 2 (7B) achieved a performance of 41% accuracy at best, where the majority class accuracy is 41% and the average chance accuracy is 27%. These data will support development of new models and approaches for detecting (and reasoning about) motific information in text.

motif, culture, natural language, artificial intel↗

Machine learning-based prediction of enzyme substrate scope: Application to bacterial nitrilases

Predicting the range of substrates accepted by an enzyme from its amino acid sequence is challenging. Although sequenc- and structure-based annotation approaches are often accurate for predicting broad categories of substrate specificity, they generally cannot predict which specific molecules will be accepted as substrates for a given enzyme, particularly within a class of closely related molecules. Combining targeted experimental activity data with structural modeling, ligand docking, and physicochemical properties of proteins and ligands with various machine learning models provides complementary information that can lead to accurate predictions of substrate scope for related enzymes. Here we describe such an approach that can predict the substrate scope of bacterial nitrilases, which catalyze the hydrolysis of nitrile compounds to the corresponding carboxylic acids and ammonia. Each of the four machine learning models (logistic regression, random forest, gradient-boosted decision trees, and support vector machines) performed similarly (average ROC = 0.9, average accuracy = ~82%) for predicting substrate scope for this dataset, although random forest offers some advantages. Finally, this approach is intended to be highly modular with respect to physicochemical property calculations and software used for structural modeling and docking.

59 BASIC BIOLOGICAL SCIENCES↗

A Preprocessing Tool for Enhanced Ion Mobility–Mass Spectrometry-Based Omics Workflows

The ability to improve the data quality of ion mobility–mass spectrometry (IM-MS) measurements is of great importance for enabling modular and efficient computational workflows and gaining better qualitative and quantitative insights from complex biological and environmental samples. We developed the PNNL PreProcessor, a standalone and user-friendly software housing various algorithmic implementations to generate new MS-files with enhanced signal quality and in the same instrument format. Different experimental approaches are supported for IM-MS based on Drift-Tube (DT) and Structures for Lossless Ion Manipulations (SLIM), including liquid chromatography (LC) and infusion analyses. The algorithms extend the dynamic range of the detection system, while reducing file sizes for faster and memory-efficient downstream processing. Specifically, multidimensional smoothing improves peak shapes of poorly defined low-abundance signals, and saturation repair reconstructs the intensity profile of high-abundance peaks from various analyte types. Further, other functionalities are data compression and interpolation, IM demultiplexing, noise filtering by low intensity threshold and spike removal, and exporting of acquisition metadata. Several advantages of the tool are illustrated, including an increase of 19.4% in lipid annotations and a two-times faster processing of LC-DT IM-MS data-independent acquisition spectra from a complex lipid extract of a standard human plasma sample. The software is freely available at https://omics.pnl.gov/software/pnnl-preprocessor.

59 BASIC BIOLOGICAL SCIENCES↗

Standardized Residue Numbering and Secondary Structure Nomenclature in the Class D β-Lactamases

Over 1370 class D β-lactamases are currently known, and they pose a serious threat to the effective treatment of many infectious diseases, particularly in some pathogenic bacteria where evolving carbapenemase activity has been reported. Detailed understanding of their molecular biology, enzymology, and structural biology are critically important, but the lack of a standardized residue numbering scheme and inconsistent secondary structure annotation has made comparative analyses sometimes difficult and cumbersome. Compounding this, in the post-AlphaFold world where we currently find ourselves, an extraordinary wealth of detailed structural information on these enzymes is literally at our fingertips; therefore it is vitally important that a standard numbering system is in place to facilitate the accurate and straightforward analysis of their structures. In conclusion, here we present a residue numbering and secondary structure scheme for the class D enzymes based on the sequence and structure of OXA-48 and apply it to test targets to demonstrate the ease with which it can be used.

59 BASIC BIOLOGICAL SCIENCES↗

Protein remote homology detection and structural alignment using deep learning

Exploiting sequence–structure–function relationships in biotechnology requires improved methods for aligning proteins that have low sequence similarity to previously annotated proteins. We develop two deep learning methods to address this gap, TM-Vec and DeepBLAST. TM-Vec allows searching for structure–structure similarities in large sequence databases. It is trained to accurately predict TM-scores as a metric of structural similarity directly from sequence pairs without the need for intermediate computation or solution of structures. Once structurally similar proteins have been identified, DeepBLAST can structurally align proteins using only sequence information by identifying structurally homologous regions between proteins. It outperforms traditional sequence alignment methods and performs similarly to structure-based alignment methods. We show the merits of TM-Vec and DeepBLAST on a variety of datasets, including better identification of remotely homologous proteins compared with state-of-the-art sequence alignment and structure prediction methods.

59 BASIC BIOLOGICAL SCIENCES↗

Chromosome‐level Thlaspi arvense genome provides new tools for translational research and for a newly domesticated cash cover crop of the cooler climates

Summary Thlaspi arvense (field pennycress) is being domesticated as a winter annual oilseed crop capable of improving ecosystems and intensifying agricultural productivity without increasing land use. It is a selfing diploid with a short life cycle and is amenable to genetic manipulations, making it an accessible field‐based model species for genetics and epigenetics. The availability of a high‐quality reference genome is vital for understanding pennycress physiology and for clarifying its evolutionary history within the Brassicaceae. Here, we present a chromosome‐level genome assembly of var. MN106‐Ref with improved gene annotation and use it to investigate gene structure differences between two accessions (MN108 and Spring32‐10) that are highly amenable to genetic transformation. We describe non‐coding RNAs, pseudogenes and transposable elements, and highlight tissue‐specific expression and methylation patterns. Resequencing of forty wild accessions provided insights into genome‐wide genetic variation, and QTL regions were identified for a seedling colour phenotype. Altogether, these data will serve as a tool for pennycress improvement in general and for translational research across the Brassicaceae.

59 BASIC BIOLOGICAL SCIENCES↗

Functional Analysis of H + -Pumping Membrane-Bound Pyrophosphatase, ADP-Glucose Synthase, and Pyruvate Phosphate Dikinase as Pyrophosphate Sources in Clostridium thermocellum

Increased understanding of the central metabolism of C. thermocellum is important from a fundamental as well as from a sustainability and industrial perspective. In addition to showing that H + -pumping membrane-bound PPase, glycogen cycling, a Ppdk–malate shunt cycle, and acetate cycling are not significant sources of PP i supply, this study adds functional annotation of four genes and availability of an updated PP i stoichiometry from biosynthesis to the scientific domain.

Clostridium thermocellum↗

Genome-Scale Transcription-Translation Mapping Reveals Features of Zymomonas mobilis Transcription Units and Promoters

Efforts to rationally engineer synthetic pathways in Zymomonas mobilis are impeded by a lack of knowledge and tools for predictable and quantitative programming of gene regulation at the transcriptional, posttranscriptional, and posttranslational levels. With the detailed functional characterization of the Z. mobilis genome presented in this work, we provide crucial knowledge for the development of synthetic genetic parts tailored to Z. mobilis . This information is vital as researchers continue to develop Z. mobilis for synthetic biology applications. Our methods and statistical analyses also provide ways to rapidly advance the understanding of poorly characterized bacteria via empirical data that enable the experimental validation of sequence-based prediction for genome characterization and annotation.

59 BASIC BIOLOGICAL SCIENCES↗