Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “text extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

KAM (Knowledge Acquisition Module): A tool to simplify the knowledge acquisition process

Analysts, knowledge engineers and information specialists are faced with increasing volumes of time-sensitive data in text form, either as free text or highly structured text records. Rapid access to the relevant data in these sources is essential. However, due to the volume and organization of the contents, and limitations of human memory and association, frequently: (1) important information is not located in time; (2) reams of irrelevant data are searched; and (3) interesting or critical associations are missed due to physical or temporal gaps involved in working with large files. The Knowledge Acquisition Module (KAM) is a microcomputer-based expert system designed to assist knowledge engineers, analysts, and other specialists in extracting useful knowledge from large volumes of digitized text and text-based files. KAM formulates non-explicit, ambiguous, or vague relations, rules, and facts into a manageable and consistent formal code. A library of system rules or heuristics is maintained to control the extraction of rules, relations, assertions, and other patterns from the text. These heuristics can be added, deleted or customized by the user. The user can further control the extraction process with optional topic specifications. This allows the user to cluster extracts based on specific topics. Because KAM formalizes diverse knowledge, it can be used by a variety of expert systems and automated reasoning applications. KAM can also perform important roles in computer-assisted training and skill development. Current research efforts include the applicability of neural networks to aid in the extraction process and the conversion of these extracts into standard formats.

Gettig, Gary A.↗

NASA Taxonomies for Searching Problem Reports and FMEAs

Many types of hazard and risk analyses are used during the life cycle of complex systems, including Failure Modes and Effects Analysis (FMEA), Hazard Analysis, Fault Tree and Event Tree Analysis, Probabilistic Risk Assessment, Reliability Analysis and analysis of Problem Reporting and Corrective Action (PRACA) databases. The success of these methods depends on the availability of input data and the analysts knowledge. Standard nomenclature can increase the reusability of hazard, risk and problem data. When nomenclature in the source texts is not standard, taxonomies with mapping words (sets of rough synonyms) can be combined with semantic search to identify items and tag them with metadata based on a rich standard nomenclature. Semantic search uses word meanings in the context of parsed phrases to find matches. The NASA taxonomies provide the word meanings. Spacecraft taxonomies and ontologies (generalization hierarchies with attributes and relationships, based on terms meanings) are being developed for types of subsystems, functions, entities, hazards and failures. The ontologies are broad and general, covering hardware, software and human systems. Semantic search of Space Station texts was used to validate and extend the taxonomies. The taxonomies have also been used to extract system connectivity (interaction) models and functions from requirements text. Now the Reconciler semantic search tool and the taxonomies are being applied to improve search in the Space Shuttle PRACA database, to discover recurring patterns of failure. Usual methods of string search and keyword search fall short because the entries are terse and have numerous shortcuts (irregular abbreviations, nonstandard acronyms, cryptic codes) and modifier words cannot be used in sentence context to refine the search. The limited and fixed FMEA categories associated with the entries do not make the fine distinctions needed in the search. The approach assigns PRACA report titles to problem classes in the taxonomy. Each ontology class includes mapping words - near-synonyms naming different manifestations of that problem class. The mapping words for Problems, Entities and Functions are converted to a canonical form plus any of a small set of modifier words (e.g. non-uniformity NOT + UNIFORM.) The report titles are parsed as sentences if possible, or treated as a flat sequence of word tokens if parsing fails. When canonical forms in the title match mapping words, the PRACA entry is associated with the corresponding Problem, Entity or Function in the ontology. The user can search for types of failures associated with types of equipment, clustering by type of problem (e.g., all bearings found with problems of being uneven: rough, irregular, gritty ). The results could also be used for tagging PRACA report entries with rich metadata. This approach could also be applied to searching and tagging failure modes, failure effects and mitigations in FMEAs. In the pilot work, parsing 52K+ truncated titles (the test cases that were available), has resulted in identification of both a type of equipment and type of problem in about 75% of the cases. The results are displayed in a manner analogous to Google search results. The effort has also led to the enrichment of the taxonomy, adding some new categories and many new mapping words. Further work would make enhancements that have been identified for improving the clustering and further reducing the false alarm rate. (In searching for recurring problems, good clustering is more important than reducing false alarms). Searching complete PRACA reports should lead to immediate improvement.

Malin, Jane T.↗

Algorithms for software development version control and change detection

Simple computer algorithms for processing source program and text data files in order to extract change detection, version control, version history, and current status information are described. These algorithms presuppose that it is possible to attach to each record of the source files a 6-character code, placed within delimiters that will cause the compiler, or other using program, to ignore this code field. The code contains a 2-character code for a character-by-character position-sensitive checksum of the record, another for the record number in the file, and a third for the data on which the encoding took place. Once the source file has been thus encoded, it is possible to detect the following transactions on the file since the most recent version coding; (1) addition of new records (having no version code), (2) modification of existing records, (3) deletion of a number of records, (4) movement and/or duplication of existing records, and (5) modification and duplication of records. In addition, it is possible to extract a version history of the number of records created or modified by date. A special file listing program is described which prints the file records without showing the version codes, but places a "change bar" at the right margin whenever a change is detected. The program also provides a list of changed pages and a version history.

Tausworthe, R. C.↗

DOCU-TEXT: A tool before the data dictionary

DOCU-TEXT, a proprietary software package that aids in the production of documentation for a data processing organization and can be installed and operated only on IBM computers is discussed. In organizing information that ultimately will reside in a data dictionary, DOCU-TEXT proved to be a useful documentation tool in extracting information from existing production jobs, procedure libraries, system catalogs, control data sets and related files. DOCU-TEXT reads these files to derive data that is useful at the system level. The output of DOCU-TEXT is a series of user selectable reports. These reports can reflect the interactions within a single job stream, a complete system, or all the systems in an installation. Any single report, or group of reports, can be generated in an independent documentation pass.

Carter, B.↗

Database tomography for commercial application

Database tomography is a method for extracting themes and their relationships from text. The algorithms, employed begin with word frequency and word proximity analysis and build upon these results. When the word 'database' is used, think of medical or police records, patents, journals, or papers, etc. (any text information that can be computer stored). Database tomography features a full text, user interactive technique enabling the user to identify areas of interest, establish relationships, and map trends for a deeper understanding of an area of interest. Database tomography concepts and applications have been reported in journals and presented at conferences. One important feature of the database tomography algorithm is that it can be used on a database of any size, and will facilitate the users ability to understand the volume of content therein. While employing the process to identify research opportunities it became obvious that this promising technology has potential applications for business, science, engineering, law, and academe. Examples include evaluating marketing trends, strategies, relationships and associations. Also, the database tomography process would be a powerful component in the area of competitive intelligence, national security intelligence and patent analysis. User interests and involvement cannot be overemphasized.

Kostoff, Ronald N.↗

Solar Astronomy Data Base: Packaged Information on Diskette

In its role as a library, the National Geophysical Data Center has transferred to diskette a collection of small, digital files of routinely measured solar indices for use on an IBM-compatible desktop computer. Recording these observations on diskette allows the distribution of specialized information to researchers with a wide range of expertise in computer science and solar astronomy. Every data set was made self-contained by including formats, extraction utilities, and plain-language descriptive text. Moreover, for several archives, two versions of the observations are provided - one suitable for display, the other for analysis with popular software packages. Since the files contain no control characters, each one can be modified with any text editor.

Mckinnon, John A.↗

Event Report for The Ethical Artificial Intelligence Quantification Workshop

Artificial Intelligence (AI) is a powerful emerging technology area which requires special attention to using it ethically. AI ethics is still an emerging field, and the partners for this workshop and report seek to move AI ethics discussion ahead by experimenting with ways to measure AI ethics criteria. The following document describes the outcomes and learnings from The Ethical Artificial Intelligence Quantification Workshop held at the National Institute for Aerospace (NIA), Hampton, Virginia on May 12th, 2022. The purpose of the workshop was for participants to evaluate and experiment-with the methodology and process presented by AIEthics.World in cooperation with Intel Corporation. The meeting participants learned about the Ethical AI Certification and Maturity Model™ and applied the methodology to selected notional AI systems. The workshop facilitated the evaluation of the maturity of the AI system according to ethical considerations relevant to NASA, NIA and other participants. The workshop consisted of three main phases. The first phase focused on understanding and summarizing NASA’s ethical approaches, mission and values based on published documentation, discussions and individual insights & opinions of participants. This information was prioritized, weighted, ordered, and quantified in phase two, to formulate an alignment between human values (ethics) and their applicability to AI systems during all lifecycle phases. The first two phases were summarized as a form of ethical genealogy for artificial intelligence, specific to NASA’s ethical approaches. In the third and last phase of the workshop the participants evaluated notional examples of artificial intelligence to qualify and quantify its ability to adhere to the organizational ethics approaches, using the Ethical AI Certification and Maturity Model™. The workshop uses the concept of genealogy, in the traditional sense: the study and traceability of lines of ancestors in the process of evolutionary development from earlier forms. However, as it is applied to an Ethical AI definition, it is providing the insights to the necessary and mandatory traceability of content, data, metrics, telemetry, elements, and structures which are used in the AI’s lifecycle to foster and measure AI ethics in all steps of its lifecycle. The Ethical Artificial Intelligence Quantification Workshop provided NASA with the opportunity to apply the Ethical AI Certification and Maturity Model™, in combination with existing and well-known decision-making and quality control methods to identify the metrics and measurements for an Ethical AI and assess its ethical condition and quality aligned with NASA ethics approaches. The result of the workshop is the capacity for NASA to apply the maturity model assessment to its AI Systems as desired and if necessary, publish the ability of these AI Systems to adhere to the organizational ethical goals. AI ethics frameworks need to be customized for each application domain, for example, individual NASA Mission Directorates. General principles that work in one area such as AI/Machine Learning-based text analysis (the ethics of information-extraction) may need to be adapted for another such as sense-and-avoid decision-making in a flight environment. The workshop was conducted among approximately twenty NASA subject matter experts, so the elements noted above should be considered examples, not definitive NASA ethical AI principles, genealogy, etc. Generating a definitive AI ethics framework for an organization as diverse as NASA would require far more discussion, debate, review, etc. However, the workshop provided valuable insight into mechanisms and processes for quantifying AI ethical qualities.

Artificial Intelligence↗

Digitizing Named Entities Found Within Letters of Agreement

Letters of Agreement (LOAs) are text-based air traffic control documents that contain procedures and actions agreed upon by the different parties, typically two or more FAA facilities, that are subject to an agreement. The documents contain among other things generic constraints, which are explicit and implicit combinations of procedures that limit a flight’s trajectory and affects pilot actions. For example, a controller may be required, to assign a specific altitude to an aircraft crossing the boundary between two airspaces. Although LOA generic constraints directly impact the trajectory of an aircraft, they are not currently available in a digital form that can be used for (or directly ingested into automated) flight planning. Instead, the constraints are manually input into an onboard or ground based system. LOA documents are primarily stored at a controlling facility and the generic constraints are implemented by experienced air traffic controllers and pilots primarily using voice instructions. This increases the workload of the controllers, likelihood of error (e.g., due to noisy communication) and makes it impractical for implementation with unmanned aircraft. Therefore, steps must be taken to make existing constraints machine interpretable to enable e.g., automated handoffs which in turn would reduce controller workload. With recent advances in natural language processing, especially the rise in digitization of text documents (e.g., medical documents) and automated extraction of information therein, it is now possible to extract flight specific constraints from LOAs. The goal of this work is to digitize named entities through a combination of natural language processing tasks: named entity disambiguation, toponym resolution, and numeric parsing to extract general constraint components contained within LOAs, herein referred to as Entity Enhancement (EE). Starting with a small list of named entities (e.g., ARTCC, Tower, Altitude and Speed), EE can extract the named entities while simultaneously converting the string-based output into a digital format using an ensemble of processes like rule-based gazetteers and syntactic-lexical patterns. The digital format contains a diverse set of information based on the entity label in question, ranging from standardized facility names to units of measure (e.g., feet) and other numeric information. Upon validating our approach using a truth dataset, we show an overall F1-Score of 0.71 for the extraction process. Looking beyond entity enhancement, we are also working towards the goal of completely digitizing the general constraints by performing EE and fitting them into a standardized exchange model (XM) such as the Aeronautical Information Exchange Model (AIXM). This will allow for easy distribution and dissemination of LOA constraints to air users, better searchability within documents, and enable ingestion into automated flight planning. Finally, we show a preliminary version of the proposed XM architecture and demonstrate how the model can be populated from the EE output.

Stephen S. B. Clarke↗

GMI-IPS: Python Processing Software for Aircraft Campaigns

NASA's Atmospheric Tomography Mission (ATom) seeks to understand the impact of anthropogenic air pollution on gases in the Earth's atmosphere. Four flight campaigns are being deployed on a seasonal basis to establish a continuous global-scale data set intended to improve the representation of chemically reactive gases in global atmospheric chemistry models. The Global Modeling Initiative (GMI), is creating chemical transport simulations on a global scale for each of the ATom flight campaigns. To meet the computational demands required to translate the GMI simulation data to grids associated with the flights from the ATom campaigns, the GMI ICARTT Processing Software (GMI-IPS) has been developed and is providing key functionality for data processing and analysis in this ongoing effort. The GMI-IPS is written in Python and provides computational kernels for data interpolation and visualization tasks on GMI simulation data. A key feature of the GMI-IPS, is its ability to read ICARTT files, a text-based file format for airborne instrument data, and extract the required flight information that defines regional and temporal grid parameters associated with an ATom flight. Perhaps most importantly, the GMI-IPS creates ICARTT files containing GMI simulated data, which are used in collaboration with ATom instrument teams and other modeling groups. The initial main task of the GMI-IPS is to interpolate GMI model data to the finer temporal resolution (1-10 seconds) of a given flight. The model data includes basic fields such as temperature and pressure, but the main focus of this effort is to provide species concentrations of chemical gases for ATom flights. The software, which uses parallel computation techniques for data intensive tasks, linearly interpolates each of the model fields to the time resolution of the flight. The temporally interpolated data is then saved to disk, and is used to create additional derived quantities. In order to translate the GMI model data to the spatial grid of the flight path as defined by the pressure, latitude, and longitude points at each flight time record, a weighted average is then calculated from the nearest neighbors in two dimensions (latitude, longitude). Using SciPya's Regular Grid Interpolator, interpolation functions are generated for the GMI model grid and the calculated weighted averages. The flight path points are then extracted from the ATom ICARTT instrument file, and are sent to the multi-dimensional interpolating functions to generate GMI field quantities along the spatial path of the flight. The interpolated field quantities are then written to a ICARTT data file, which is stored for further manipulation. The GMI-IPS is aware of a generic ATom ICARTT header format, containing basic information for all flight campaigns. The GMI-IPS includes logic to edit metadata for the derived field quantities, as well as modify the generic header data such as processing dates and associated instrument files. The ICARTT interpolated data is then appended to the modified header data, and the ICARTT processing is complete for the given flight and ready for collaboration. The output ICARTT data adheres to the ICARTT file format standards V1.1. The visualization component of the GMI-IPS uses Matplotlib extensively and has several functions ranging in complexity. First, it creates a model background curtain for the flight (time versus model eta levels) with the interpolated flight data superimposed on the curtain. Secondly, it creates a time-series plot of the interpolated flight data. Lastly, the visualization component creates averaged 2D model slices (longitude versus latitude) with overlaid flight track circles at key pressure levels. The GMI-IPS consists of a handful of classes and supporting functionality that have been generalized to be compatible with any ICARTT file that adheres to the base class definition. The base class represents a generic ICARTT entry, only defining a single time entry and 3D spatial positioning parameters. Other classes inherit from this base class; several classes for input ICARTT instrument files, which contain the necessary flight positioning information as a basis for data processing, as well as other classes for output ICARTT files, which contain the interpolated model data. Utility classes provide functionality for routine procedures such as: comparing field names among ICARTT files, reading ICARTT entries from a data file and storing them in data structures, and returning a reduced spatial grid based on a collection of ICARTT entries. Although the GMI-IPS is compatible with GMI model data, it can be adapted with reasonable effort for any simulation that creates Hierarchical Data Format (HDF) files. The same can be said of its adaptability to ICARTT files outside of the context of the ATom mission. The GMI-IPS contains just under 30,000 lines of code, eight classes, and a dozen drivers and utility programs. It is maintained with GIT source code management and has been used to deliver processed GMI model data for the ATom campaigns that have taken place to date.

Damon, M. R.↗

A transportable system for management and exchange of programs and other text

A system is presented for exchanging information on tape, that allows information about the data to be included with the data. The system is designed for portability, but requires a few simple machine dependent modules. These modules are available for a variety of machines, and a bootstrapping procedure is provided. The system allows content selected reading of the tape, and a simple text editing facility is provided. Although the system recognizes 30 commands, information may be extracted from the tape by using as few as three commands. In addition to its use for information exchange, the system is expected to find use in maintaining large libraries of text.

Snyder, W. V.↗

What Went Wrong: A Survey of Wildfire UAS Mishaps through Named Entity Recognition

Increasingly, unmanned aircraft systems (UAS) are being applied to wildfire incidents for tasks such as mapping, aerial ignition, and delivery. As a result, aviation incident reporting systems for wildfires are beginning to accumulate data related to UAS mishaps in wildfire response. In this research, we apply state-of-the-art natural language processing (NLP) techniques to develop a custom Named Entity Recognition (NER) model which extracts entities relevant to safety analysts. The custom NER model is built by fine-tuning an existing Bidirectional Encoder Representations from Transformers (BERT) model, resulting in a generalizable NER model that can extract engineering relevant entities including failure modes, causes, effects, control processes, and recommendations from failure-relevant text. This model performs passably, with a weighted average f1 score of 0.33 across entity types, indicating more labeled training data is needed. Extracted entities are used to form a Failure Modes and Effects Analysis (FMEA)-style survey of wildfire UAS mishaps reported using the SAFECOM system. Similar mishaps are manually clustered and reported as single rows within an FMEA. Foreach cluster, we compute frequency, severity, and overall riskin accordance with FAA standards. This methodology can beapplied as part of a broader safety management system totrack trends in mishaps (e.g., likelihood, severity) and discoverknowledge (e.g., causes, effects) that can be utilized to improvesafety outcomes and system performance.

Machine Learning↗

BOREAS TE-2 NSA Soil Lab Data

This data set contains the major soil properties of soil samples collected in 1994 at the tower flux sites in the Northern Study Area (NSA). The soil samples were collected by Hugo Veldhuis and his staff from the University of Manitoba. The mineral soil samples were largely analyzed by Barry Goetz, under the supervision of Dr. Harold Rostad at the University of Saskatchewan. The organic soil samples were largely analyzed by Peter Haluschak, under the supervision of Hugo Veldhuis at the Centre for Land and Biological Resources Research in Winnipeg, Manitoba. During the course of field investigation and mapping, selected surface and subsurface soil samples were collected for laboratory analysis. These samples were used as benchmark references for specific soil attributes in general soil characterization. Detailed soil sampling, description, and laboratory analysis were performed on selected modal soils to provide examples of common soil physical and chemical characteristics in the study area. The soil properties that were determined include soil horizon; dry soil color; pH; bulk density; total, organic, and inorganic carbon; electric conductivity; cation exchange capacity; exchangeable sodium, potassium, calcium, magnesium, and hydrogen; water content at 0.01, 0.033, and 1.5 MPascals; nitrogen; phosphorus: particle size distribution; texture; pH of the mineral soil and of the organic soil; extractable acid; and sulfur. These data are stored in ASCII text files. The data files are available on a CD-ROM (see document number 20010000884), or from the Oak Ridge National Laboratory (ORNL) Distributed Active Archive Center (DAAC).

Veldhuis, Hugo↗

BOREAS TE-1 SSA Soil Lab Data

This data set was collected by TE-1 to provide a set of soil properties for BOREAS investigators in the SSA. The soil samples were collected at sets of soil pits in 1993 and 1994. Each set of soil pits was in the vicinity of one of the five flux towers in the BOREAS SSA. The collected soil samples were sent to a lab, where the major soil properties were determined. These properties include, but are not limited to, soil horizon; dry soil color; pH; bulk density; total, organic, and inorganic carbon; electric conductivity; cation exchange capacity; exchangeable sodium, potassium, calcium, magnesium, and hydrogen; water content at 0.01, 0.033, and 1.5 MPascals; nitrogen; phosphorus; particle size distribution; texture; pH of the mineral soil and of the organic soil; extractable acid; and sulfur. The data are stored in tabular ASCII text files. The data files are available on a CD-ROM (see document number 20010000884), or from the Oak Ridge National Laboratory (ORNL) Distributed Active Archive Center (DAAC).

Hall, Forrest G.↗

Towards Automated Analytics of Research Publications

For readers of scientific publications it remains a big challenge to unambiguously relate the published research with the data used. To a substantial degree it is attributed to authors, journals, editors, and reviewers not prioritizing correct data citation, which impacts traceability, repeatability, and giving credits to published authors and their funding sources. Furthermore, uniform classification of the content of the published research is hampered by journals using journal specific topics and letting authors to assign free text keywords to their papers. We demonstrate automated analytics methods for extracting and relating datasets used and the research application areas by processing 1,300 research papers that referenced the NASA Giovanni service (but probably not the datasets in particular) as supporting their publication process. This presentation was given during the 2022 ESIP January meeting held virtually in January 2022.

Irina Gerasimov↗

Semantic Theme Analysis of Pilot Incident Reports

Pilots report accidents or incidents during take-off, on flight and landing to airline authorities and Federal aviation authority as well. The description of pilot reports for an incident contains technical terms related to Flight instruments and operations. Normal text mining approaches collect keywords from text documents and relate them among documents that are stored in database. Present approach will extract specific theme analysis of incident reports and semantically relate hierarchy of terms assigning weights of themes. Once the theme extraction has been performed for a given document, a unique key can be assigned to that document to cross linking the documents. Semantic linking will be used to categorize the documents based on specific rules that can help an end-user to analyze certain types of accidents. This presentation outlines the architecture of text mining for pilot incident reports for autonomous categorization of pilot incident reports using semantic theme analysis.

Thirumalainambi, Rajkumar↗

Topic Modeling of NASA Space System Problem Reports: Research in Practice

Problem reports at NASA are similar to bug reports: they capture defects found during test, post-launch operational anomalies, and document the investigation and corrective action of the issue. These artifacts are a rich source of lessons learned for NASA, but are expensive to analyze since problem reports are comprised primarily of natural language text. We apply topic modeling to a corpus of NASA problem reports to extract trends in testing and operational failures. We collected 16,669 problem reports from six NASA space flight missions and applied Latent Dirichlet Allocation topic modeling to the document corpus. We analyze the most popular topics within and across missions, and how popular topics changed over the lifetime of a mission. We find that hardware material and flight software issues are common during the integration and testing phase, while ground station software and equipment issues are more common during the operations phase. We identify a number of challenges in topic modeling for trend analysis: 1) that the process of selecting the topic modeling parameters lacks definitive guidance, 2) defining semantically-meaningful topic labels requires nontrivial effort and domain expertise, 3) topic models derived from the combined corpus of the six missions were biased toward the larger missions, and 4) topics must be semantically distinct as well as cohesive to be useful. Nonetheless,topic modeling can identify problem themes within missions and across mission lifetimes, providing useful feedback to engineers and project managers.

Data Mining↗

The flat fielding and achievable signal-to-noise of the MAMA detectors

The Space Telescope Imaging Spectrograph (STIS) was designed to achieve a signal-to-noise (S/N) of at least 100:1 per resolution element. Multi-Anode Microchannel Arrays (MAMA) observations during Servicing Mission Orbital Verification (SMOV) confirm that this specification can be met. From analysis of a single spectrum of GD153, with counting statistics of approximately 165 a S/N of approximately 125 is achieved per spectral resolution element in the far ultraviolet (FUV) over the spectral range of 1280A to 1455A. Co-adding spectra of GRW+7OD5824 to increase the counting statistics to approximately 300 yields a S/N of approximately 190 per spectral resolution element over the region extending from 1347A to 1480A in the FUV. In the near ultraviolet (NUV), a single spectrum of GRW+7OD5824 with counting statistics of approximately 200 yields a S/N of approximately 150 per spectral resolution element over the spectral region extending from 2167 to 2520A. Details of the flat field construction, the spectral extraction, and the definition of a spectral resolution element will be described in the text.

Kaiser, Mary Elizabeth↗

Creative Analytics of Mission Ops Event Messages

Historically, tremendous effort has been put into processing and displaying mission health and safety telemetry data; and relatively little attention has been paid to extracting information from missions time-tagged event log messages. Todays missions may log tens of thousands of messages per day and the numbers are expected to dramatically increase as satellite fleets and constellations are launched, as security monitoring continues to evolve, and as the overall complexity of ground system operations increases. The logs may contain information about orbital events, scheduled and actual observations, device status and anomalies, when operators were logged on, when commands were resent, when there were data drop outs or system failures, and much much more. When dealing with distributed space missions or operational fleets, it becomes even more important to systematically analyze this data. Several advanced information systems technologies make it appropriate to now develop analytic capabilities which can increase mission situational awareness, reduce mission risk, enable better event-driven automation and cross-mission collaborations, and lead to improved operations strategies: Industry Standard for Log Messages. The Object Management Group (OMG) Space Domain Task Force (SDTF) standards organization is in the process of creating a formal standard for industry for event log messages. The format is based on work at NASA GSFC. Open System Architectures. The DoD, NASA, and others are moving towards common open system architectures for mission ground data systems based on work at NASA GSFC with the full support of the commercial product industry and major integration contractors. Text Analytics. A specific area of data analytics which applies statistical, linguistic, and structural techniques to extract and classify information from textual sources. This presentation describes work now underway at NASA to increase situational awareness through the collection of non-telemetry mission operations information into a common log format and then providing display and analytics tools to provide in-depth assessment of the log contents. The work includes: Common interface formats for acquiring time-tagged text messages Conversion of common files for schedules, orbital events, and stored commands to the common log format Innovative displays to depict thousands of messages on a single display Structured English text queries against the log message data store, extensible to a more mature natural language query capability Goal of speech-to-text and text-to-speech additions to create a personal mission operations assistant to aid on-console operations. A wide variety of planned uses identified by the mission operations teams will be discussed.

events↗