Natural Language Processing for Text Based Event Extraction: Identifying Events of Interest Related to Worldwide State-Sponsored Civil Nuclear Power
Beginning in FY20, SRNL was funded by the National Nuclear Security Administration’s Office of Defense Nuclear Non-Proliferation Research and Development to develop a prototype natural language processing/natural language understating machine learning-based modeling and analysis pipeline to extract and forecast events of interest from massive open data sources. The working hypothesis within the approach is that contextual shifts in key words and phrases act as indicators of events of interest over time. Therefore, by identifying points in time where contextual shifts occur, events of interest can be extracted along with explicit and implicit connections of entities and activities. The development of the preliminary prototype pipeline proved successful, meriting further testing of the pipeline on more broad topical domains and in a worldwide data environment. Therefore, SRNL, in collaboration with the Sanghani Center for Artificial Intelligence and Data Analytics at Virginia Tech, have continued development with a test case of identifying events of interest related to worldwide state-sponsored civil nuclear power in open data sources. In the first year of this follow-on effort, the team has curated domain-specific data corpuses using an automated scheme and applied the modeling and analysis pipeline. This robust, focused, and efficient approach consists of an ensemble of analyses applied to time dependent word embedding models that are trained on the data corpuses. In this report, the team has demonstrated the capability of the existing pipeline (as development has continued in parallel) by exploring several specific case-studies centered around Rosatom’s international activities regarding the planning, construction, operation, and/or shutdown of nuclear reactors. A basic timeline events has been generated by manually cataloging known “milestone” events that have occurred at reactors in Turkey, Finland, Hungary, and Egypt and compared with the output of the modeling pipeline. In this approach, the team has characterized the lead time using the prototype pipeline, as well as the ability to capture relevant information, which proved 100% successful. A deep dive example of the Akkuyu reactor (Turkey) is presented that shows the breadth of information that can be captured using the approach. In this case study, events were extracted pertaining to the planning/construction of Akkuyu including protests from the population, information campaigns in response to the protests, forged regulatory documents and lawsuits, budgetary/shareholder information, geopolitical tensions, and the various construction milestones. This has demonstrated the pipeline’s utility as a research aid or real-time event extraction tool, where summary-level information and detailed text extractions from millions of articles or Tweets across long time periods can be generated with significantly less effort than current techniques.