Dynamic Trust and Authority Assignment in Autonomous Multiagent Teams Managing Uncertainty in Decision-making
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
A major weakness of a Large Language Model (LLM) is its tendency to accept information at face value, often leading to injection of erroneous information and inducing a greater probability of hallucinating non-existent information. While Retrieval Augmented Generation (RAG) uses external knowledge sources to bolster LLMs through grounded truth, this work seeks to explore methods to engender a LLM with an intrinsic capability to evaluate an input’s believability without relying on external knowledge sources. We investigate unifying a LLM with a Knowledge Graph (KG) and using the KG to reinforce the LLM’s internal word embedding while also maintaining belief metrics along the edge’s in the KG.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
In essence, project management is about people. Virtually every successful project is defined by good relations between the people involved. In the same way, nearly every failed or troubled project is about poor relationships between the people involved. Let's consider one type of relationship: the one between the government and the contractor. It's easy to say that a contractor must earn the government's trust, but what does that mean in practice? Who needs to earn whose trust? What's the timeline for doing that? How does anyone know when he or she is trusted? What is the relationship supposed to be like before one feels like trust has really been established? So many questions it makes my head hurt. I have always found it better to begin a relationship assuming that everyone is trustworthy until, and unless, something occurs to belie trust.
The Space Human Factors Engineering (SHFE) Standing Review Panel (SRP) evaluated 22 gaps and 39 tasks in the three risk areas assigned to the SHFE Project. The area where tasks were best designed to close the gaps and the fewest gaps were left out was the Risk of Reduced Safety and Efficiency dire to Inadequate Design of Vehicle, Environment, Tools or Equipment. The areas where there were more issues with gaps and tasks, including poor or inadequate fit of tasks to gaps and missing gaps, were Risk of Errors due to Poor Task Design and Risk of Error due to Inadequate Information. One risk, the Risk of Errors due to Inappropriate Levels of Trust in Automation, should be added. If astronauts trust automation too much in areas where it should not be trusted, but rather tempered with human judgment and decision making, they will incur errors. Conversely, if they do not trust automation when it should be trusted, as in cases where it can sense aspects of the environment such as radiation levels or distances in space, they will also incur errors. This will be a larger risk when astronauts are less able to rely on human mission control experts and are out of touch, far away, and on their own. The SRP also identified 11 new gaps and five new tasks. Although the SRP had an extremely large quantity of reading material prior to and during the meeting, we still did not feel we had an overview of the activities and tasks the astronauts would be performing in exploration missions. Without a detailed task analysis and taxonomy of activities the humans would be engaged in, we felt it was impossible to know whether the gaps and tasks were really sufficient to insure human safety, performance, and comfort in the exploration missions. The SRP had difficulty evaluating many of the gaps and tasks that were not as quantitative as those related to concrete physical danger such as excessive noise and vibration. Often the research tasks for cognitive risks that accompany poor task or information design addressed only part, but not all, of the gaps they were programmed to fill. In fact the tasks outlined will not close the gap but only scratch the surface in many cases. In other cases, the gap was written too broadly, and really should be restated in a more constrained way that can be addressed by a well-organized and complementary set of tasks. In many cases, the research results should be turned into guidelines for design. However, it was not clear whether the researchers or another group would construct and deliver these guidelines.
The FAIR principle (findable, accessible, interoperable, and reusable) governs the storage and sharing of NASA space biology and health data[1]. These guiding principles maximize reuse of data and the reproducibility of scientific findings. The NASA Open Science Data Repository (OSDR; an expansion of NASA GeneLab) was built on the FAIR principles and houses over 500 studies and close to 1000 datasets from decades of space life sciences experiments. OSDR embodies the FAIR principles through data governance that includes mediated, embargoed, and fully open access data. The FAIR data governance principles were recently proposed to be expanded to encompass a FAIREST framework for assessing research data repositories (FAIR + Engagement, Social connections, and Trust)[2]. FAIREST emphasizes the importance of data repositories engaging with the scientific community and gaining the trust of researchers regarding data quality. Trust also refers to the TRUST principles developed for assessment of digital repositories: Transparency, Responsibility, User Focus, Sustainability, Technology[3]. We present the “Open Science for Life in Space” Analysis Working Groups (AWGs) as evidence regarding the power of engagement, social connections, and trust which has enhanced OSDR’s capabilities and productivity. AWG members engage in two main activities. One, members provide feedback on OSDR scientific standards for data ingestion, curation, and reuse (study, subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability). Two, AWG members collaborate to mine-reuse OSDR data to conduct scientific analysis. With nearly 800 active members, the AWGs have resulted in 32 publications re-using OSDR data and contributed many papers in two major special issues in Cell (2020) and Nature (2024). AWGs also serve as networking groups, facilitate social connections between researchers at all levels of experience, and also have a social online ‘Forum’ used to keep members informed on projects and opportunities. This community-centric, productive, and trustworthy data culture has resulted in a broader effect with international space agencies, academics, and the commercial space sector wanting to submit their data to OSDR. Ten studies of Inspiration 4 data were recently publicly released by OSDR, as were some JAXA human data. Coming up soon in OSDR are data submissions from the European Space Agency, Virgin Galactic PIs, and SpaceX Polaris Dawn. A major benefit of OSDR is the array of standardized and uniformly formatted data (which was developed through AWG member consensus), from which visualization tools, analysis tools, and machine learning models can be built or trained. This talk will cover the Multi-Study Visualization Tool, the Environmental Data Application, RadLab, and a UCSF-NSF funded knowledge graph biomedical health discovery tool ‘SPOKE’ currently being integrated with OSDR. OSDR also provides training programs in bioinformatics and machine learning to improve the scientific community’s awareness of data availability and to boost their ability to perform data analysis. The increasing engagement of the scientific community and the public with technologies powered by artificial intelligence (AI) heightens the need for data analysis to be transparent. The AI for Life in Space initiative leverages the data products provided in OSDR to train AI models, with an emphasis on explainable and trustworthy AI, which would not be possible without FAIR data and metadata. Overall, here we will demonstrate the importance for NASA life sciences data repositories to adhere to the FAIREST framework, by providing examples and success stories from different aspects of OSDR.
Uncrewed aerial systems (UAS) show promise in urban air transport, package delivery, and emergency services. UAS efficiency can be significantly improved by having fewer operators (m) manage a greater number of vehicles (N), or the m:N architecture of operation. The current study investigates how workload affects operators’ task-allocation decision-making and potential effects of two crucial human factors: trust and self-confidence. In the context of a simulated UAS package-delivery task, 10 participants with expertise in UAS operation were recruited. Each participant reported their preferred task-allocation strategy for a set of five subtasks while watching two sets of videos with different workload levels. Perceived workload, trust, and self-confidence were also measured after each video session. Overall, participants indicated a preference for automation for most of the subtasks under the delivery mission. Trust, rather than workload and self-confidence, played a significant role in experts’ decisions of task-allocation and assignment methods. Higher trust led to higher preference for automation.
Many measurements were taken by test engineers from Hamilton Sundstrand, the prime contractor for the current EVA suit. Because the raw measurements needed to be converted to torques and combined into a final score, it was impossible to keep track of who was ahead in this phase. The final comfort and dexterity test was performed in a depressurized glove box to simulate real on-orbit conditions. Each competitor was required to exercise the glove through a defined set of finger, thumb, and wrist motions without any sign of abrasion or bruising of the competitor's hand. I learned a lot about arm fatigue! This was a pass-fail event, and both of the remaining competitors came through intact. After taking what seemed like an eternity to tally the final scores, the judges announced that I had won the competition. My glove was the only one to have achieved lower finger-bending torques than the Phase VI glove. Looking back, I see three sources of the success of this project that I believe also operate in other programs where small teams have broken new ground in aerospace technologies. These are awareness, failure, and trust. By remaining aware of the big picture, continuously asking myself, "Am I converging on a solution?" and "Am I converging fast enough?" I was able to see that my original design was not going to succeed, leading to the decision to start over. I was also aware that, had I lingered over this choice or taken time to analyze it, I would not have been ready on the first day of competition. Failure forced me to look outside conventional thinking and opened the door to innovation. Choosing to make incremental failures enabled me to rapidly climb the learning curve. Trusting my "gut" feelings-which are really an internalized accumulation of experiences-and my newly acquired skills allowed me to devise new technologies rapidly and complete both gloves just in time. Awareness, failure, and trust are intertwined: failure provides experiences that inform awareness and provide decision-making opportunities that build trust among team members and managers while opening minds to new pathways for development. All three are necessary for teams-large or small-to achieve big innovation.
This NASA conference publication contains the proceedings of the Third International Workshop on Proof-Carrying Code and Software Certification, held as part of LICS in Los Angeles, CA, USA, on August 15, 2009. Software certification demonstrates the reliability, safety, or security of software systems in such a way that it can be checked by an independent authority with minimal trust in the techniques and tools used in the certification process itself. It can build on existing validation and verification (V&V) techniques but introduces the notion of explicit software certificates, Vvilich contain all the information necessary for an independent assessment of the demonstrated properties. One such example is proof-carrying code (PCC) which is an important and distinctive approach to enhancing trust in programs. It provides a practical framework for independent assurance of program behavior; especially where source code is not available, or the code author and user are unknown to each other. The workshop wiII address theoretical foundations of logic-based software certification as well as practical examples and work on alternative application domains. Here "certificate" is construed broadly, to include not just mathematical derivations and proofs but also safety and assurance cases, or any fonnal evidence that supports the semantic analysis of programs: that is, evidence about an intrinsic property of code and its behaviour that can be independently checked by any user, intermediary, or third party. These guarantees mean that software certificates raise trust in the code itself, distinct from and complementary to any existing trust in the creator of the code, the process used to produce it, or its distributor. In addition to the contributed talks, the workshop featured two invited talks, by Kelly Hayhurst and Andrew Appel. The PCC 2009 website can be found at http://ti.arc.nasa.gov /event/pcc 091.
f autonomous systems using trust and trustworthiness is the focus of Autonomy Teaming and TRAjectories for Complex Trusted Operational Reliability (ATTRACTOR), a new NASA Convergent Aeronautical Solutions (CAS) Project. One critical research element of ATTRACTOR is explainability of the decision-making across relevant subsystems of an autonomous system. The ability to explain why an autonomous system makes a decision is needed to establish a basis of trustworthiness to safely complete a mission. Convolutional Neural Networks (CNNs) are popular visual object classifiers that have achieved high levels of classification performances without clear insight into the mechanisms of the internal layers and features. To explore the explainability of the internal components of CNNs, we reviewed three feature visualization methods in a layer-by-layer approach using aviation related images as inputs. Our approach to this is to analyze the key components of a classification event in order to generate component labels for features of the classified image at different layers of depths. For example, an airplane has wings, engines, and landing gear. These could possibly be identified somewhere in the hidden layers from the classification and these descriptive labels could be provided to a human or machine teammate while conducting a shared mission and to engender trust. Each descriptive feature may also be decomposed to a combination of primitives such as shapes and lines. We expect that knowing the combination of shapes and parts that create a classification will enable trust in the system and insight into creating better structures for the CNN.
Building a foundation for trustworthiness and trust verification in multi-asset teaming is the research challenge of Autonomy Teaming and TRAjectories for Complex Trusted Operational Reliability (ATTRACTOR). The Design Reference Mission (DRM) for ATTRACTOR is a search and rescue mission objective governed by a multi-member team consisting of human and machine operators. A crucial component to the effort is the communication between humans and autonomous agents throughout both planning and execution stages of the mission. Intuitive communication methods and modalities are posited as critical enablers for certifying trust and trustworthiness. This paper reports on the data collection and analysis conducted in support of the Human Informed Natural-language GANs Evaluation (HINGE)project to attain explainable and trusted communication between human-machine assets. Two identically curated image description datasets were acquired for HINGE, both consisting of two unique input modalities (typed vs. verbal) and retrieved in two distinct contexts (general vs. specific). The gathered datasets were assessed and compared using Parts-of-Speech (POS)features, sentence similarity metrics, and linguistic analysis. Then, the datasets were modeled and tested separately and in combination with one another using machine learning algorithms. The comparison and testing results reveal a superior dataset, by which a preferred context and input is understood, for generating image representations of missing persons using a Generative Adversarial Network (GAN).
As autonomous systems continue to grow both in use and complexity, the necessity for robust and extensible simulation-to-flight frameworks is paramount for establishing an effective architecture for autonomous systems. Hardware test flights are time-consuming and cost prohibitive during early system design and development. Simulation environments can be useful tools to accelerate algorithm development and testing. However, transitions from simulation to flight (sim-to-flight) can be challenging, unless systems are designed with this transition in mind and with the necessary capabilities built into the architecture and framework. One of the objectives of Autonomy Teaming and TRAjectories for Complex Trusted Operational Reliability (ATTRACTOR) was to design and develop a distributed mixed-reality simulation environment to begin establishing a basis for certification of autonomous systems via research into trust and trustworthiness. ATTRACTOR’s objective was to construct computational concepts of trustworthiness and justifiable trust in multi-agent autonomous teams, to inform future certification of safety-critical and time-critical autonomous systems in aviation. In this paper, we present an autonomous systems architecture and development framework paired with a persistent distributed modeling and simulation (ModSim) environment for test and evaluation of autonomous systems. They were designed under ATTRACTOR in order to measure and establish trustworthiness and trust in single-and multi-agent human-machine systems whether these machines are fixed-wing general aviation, rotary-wing Unmanned Aerial Vehicles (UAVs), ground rovers, or even spacecraft. The Autonomous Entity Operational Network (AEON) framework enables autonomous system development with an easily extensible collection of libraries and plug-n-play nodes facilitated by the Data Distribution Service (DDS) communication protocol standard. The Baseline Environment for Autonomous Modeling (BEAM) simulation environment is a distributed mixed-reality Unity™-based environment built around the same DDS communication paradigm allowing for easy integration with AEON-based autonomous applications, enabling sim-to-flight with minimal configuration changes. Using AEON and BEAM, source code that runs in simulation ports directly to hardware and has successfully flown in the National Airspace System (NAS) at NASA LaRC many times over the lifetime of ATTRACTOR.
EXECUTIVE SUMMARY The Western States Water Council (WSWC) and the NASA Western Water Applications Office (WWAO) hosted a joint workshop on technology transfer for water management in the Western U.S. The goals of the workshop were to understand how different agencies approach the technology transfer and research to operations (R2O) process, identify best practices, and discuss existing barriers to successful technology infusion into operational water resource management systems at the state and federal level. The workshop took place August 7-9, 2019 in Irvine, CA. Key outcomes of the meeting include the following:• U.S. Rep. Grace Napolitano provided opening remarks for the workshop, where she highlighted the critical value of water data and the importance of collaboration between state and federal agencies in working to advance the use of water data in water management, planning and policy. • A total of 33 participants (including remote participants) were part of the workshop. They included principal investigators and project teams supported by NASA (Cyanobacteria Assessment Network, Evapotranspiration for Western States, Evaporative Stress Index, the Airborne Snow Observatory, Satellite-based Snow Water Equivalent in the Sierra Nevadas, and Fallowed Area Mapping) as well as representatives from federal (USGS, NOAA, USBR, EPA) and state (CA, WY, OR, NE) agency partners. • One main outcome of the meeting was the consensus that successful transitions of new applications and new technologies into operations require careful planning, effective communication within and across institutions, resources and considerable time investments. In addition, there was broad agreement that significant lead time is often required to allow for identification of financial and technical resources to sustain operational use of new data, information and tools.• The meeting included remarks from U.S. Rep. Napolitano and discussions during presentations and breakout groups about key opportunities to develop best practices and streamline the technology transfer process. • For example, one key set of best practices that emerged revolved around the the importance of building trust and establishing clear lines of communication between the research and operational institutions. The conversations led to defining two key components of trust-building. The first aspect is purely technical. It requires effectively demonstrating that the proposed application meets the end user’s needs in terms of accuracy, format, resolution, latency, metadata and documentation. The second aspect of building trust involves developing sustained, productive and mutually-beneficial relationships with the partner operational agency. The best practices presented here span both the technical as well as the relational aspects of cultivating trust. • This workshop served as a first step in developing a broader community discussion around R2O in western water management. Many of the best practices and lessons learned described in this report represent starting places for action within the WWAO, WSWC and our colleagues’ institutions. • Effective implementation of the best practices that emerged from this workshop will require sustained investments of time, resources and transition planning. In recognition of this, the WSWC and the WWAO proposed continuation of discussions begun at the workshop through a series of semi-annual or annual workshops.
Trust region algorithms provide a robust iterative technique for solving non-convex unstrained optimization problems, but in many instances it is prohibitively expensive to compute high accuracy function and gradient values for the method. Of particular interest are inverse and parameter estimation problems, since function and gradient evaluations involve numerically solving large systems of differential equations. A global convergence theory is presented for trust region algorithms in which neither function nor gradient values are known exactly. The theory is formulated in a Hilbert space setting so that it can be applied to variational problems as well as the finite dimensional problems normally seen in trust region literature. The conditions concerning allowable error are remarkably relaxed: relative errors in the gradient error condition is automatically satisfied if the error is orthogonal to the gradient approximation. A technique for estimating gradient error and improving the approximation is also presented.
Viewgraphs from the Information Security and Integrity Systems seminar held at the University of Houston-Clear Lake on May 15-16, 1990 are presented. A tutorial on computer security is presented. The goals of this tutorial are the following: to review security requirements imposed by government and by common sense; to examine risk analysis methods to help keep sight of forest while in trees; to discuss the current hot topic of viruses (which will stay hot); to examine network security, now and in the next year to 30 years; to give a brief overview of encryption; to review protection methods in operating systems; to review database security problems; to review the Trusted Computer System Evaluation Criteria (Orange Book); to comment on formal verification methods; to consider new approaches (like intrusion detection and biometrics); to review the old, low tech, and still good solutions; and to give pointers to the literature and to where to get help. Other topics covered include security in software applications and development; risk management; trust: formal methods and associated techniques; secure distributed operating system and verification; trusted Ada; a conceptual model for supporting a B3+ dynamic multilevel security and integrity in the Ada runtime environment; and information intelligence sciences.
Display of information in the cockpit has long been a challenge for aircraft designers. Given the limited space in which to present information, designers have had to be extremely selective about the types and amount of flight related information to present to pilots. The general goal of cockpit display design and implementation is to ensure that displays present information that is timely, useful, and helpful. This suggests that displays should facilitate the management of perceived workload, and should allow maximal situation awareness. The formatting of current and projected weather displays represents a unique challenge. As technologies have been developed to increase the variety and capabilities of weather information available to flight crews, factors such as conflicting weather representations and increased decision importance have increased the likelihood for errors. However, if formatted optimally, it is possible that next generation weather displays could allow for clearer indications of weather trends such as developing or decaying weather patterns. Important issues to address include the integration of weather information sources, flight crew trust of displayed weather information, and the teamed reactivity of flight crews to displays of weather. Past studies of weather display reactivity and formatting have not adequately addressed these issues; in part because experimental stimuli have not approximated the complexity of modern weather displays, and in part because they have not used realistic experimental tasks or participants. The goal of the research reported here was to investigate the influence of onboard and NEXRAD agreement, range to the simulated potential weather event, and the pilot flying on flight crew deviation decisions, perceived workload, and perceived situation awareness. Fifteen pilot-copilot teams were required to fly a simulated route while reacting to weather events presented in two graphical formats on a separate visual display. Measures of flight crew reactions included performance-based measures such as deviation decision accuracy, and judgment-based measures such as perceived decision confidence, workload, situation awareness, and display trust. Results demonstrated that pilots adopted a conservative reaction strategy, often choosing to deviate from weather rather than ride through it. When onboard and NEXRAD displays did not agree, flight crews reacted in a complex manner, trusting the onboard system more but using the NEXRAD system to augment their situation awareness. Distance to weather reduced situation awareness and heightened workload levels. Overall, flight crews tended to adopt a participative leadership style marked by open communication. These results suggest that future weather displays should exploit the existing benefits of NEXRAD presentation for situation awareness while retaining the display structure and logic inherent in the onboard system.
It has been recognized that a framework based on proof-carrying code (also called semantic-based software certification in its community) could be used as a candidate software certification process for the avionics industry. To meet this goal, tools in the "trust base" of a proof-carrying code system must be qualified by regulatory authorities. A family of semantic-based software certification approaches is described, each different in expressive power, level of automation and trust base. Of particular interest is the so-called abstraction-carrying code, which can certify temporal properties. When a pure abstraction-carrying code method is used in the context of industrial software certification, the fact that the trust base includes a model checker would incur a high qualification cost. This position paper proposes a hybrid of abstraction-based and proof-based certification methods so that the model checker used by a client can be significantly simplified, thereby leading to lower cost in tool qualification.