Last semester I supervised a master's thesis project about improving the functionalities of a costumer support system with the help of semantic knowledge. One of the challenges in the project was to match categorized support tickets with un-categorized ones. Support tickets in this context are structured texts describing a problem or a question. These tickets were written by experts from the same organization who have been trained to formulate messages to individuals inside and outside the organization. I will not go into the details of the project, those who are interested can read the thesis here. What I opt to emphasis in this blog are some of the interesting outcomes of the project.
There are plenty of Natural Language Processing (NLP) resources and methods which can potentially be applied to solve the problem. When it comes to solutions that require natural language understanding, scientific approaches have previously relied on lexical semantic resources like FrameNet to capture the semantics of well-formed sentences. FrameNet encodes information about lexical units and semantic roles in a network of semantic frames, where each frame in the net represents a cognitive, real-world scenario. The computational resource, originally developed for English, has been implemented for several languages including Swedish, see Swedish FrameNet, which was developed as a part of the SweFN++ infrastructure project. However, because the language the thesis was concerned with is English, English FrameNet was the one in focus.
English FrameNet has a pre-trained model for automatically identifying and encoding words or phrases with semantic roles -- a so-called semantic role labeler. Below we see an example of a sentence taken from a support ticket, automatically processed with the FrameNet-based model. The sentence, triggered by the verb Unable, was labled with the frame Capability, the model further identified Event as the only semantic role in the sentence.
Using this model, the students analyzed 2,287 support tickets ranging from 2018-2022, with the vast majority of tickets created from 2020, as can be seen in the diagram below. At most, the model identified 3,000 frames in the last quarter (October-Decmeber) 2021. These numbers might seem overwhelming. Notwithstanding, they could be explained by the fact that one sentence can have several analyses. For example, the sentence above could also be labeled with the frame Erasing that is triggered by the verb delete.
When analyzing the data further they could identify the most frequently occuring frames. In the diagram below we see the top 7 dominating frames occuring in the data from June 2020 to June 2022, where most used frame for a quarter has rank 1.
The diagram highlights the type of support that has dominated each period. Interestingly, the frames Successful_action and Capability are among the top frames throughout the examined period. Among the top frames from 2020 we also find Using (triggered by the verbs apply, employ, operate, use, etc.) and Point_of_dispute (triggered by the nouns concern, issue, and question). Among the top frames from 2021 we also find the frame Attempt (triggered by attempt, effort, endeavor, try, etc.). Among the top frames from 2022 we also find the frames Inspecting (triggered by the verbs check, examine, inspect, etc.) and Intentionally_create (triggered by the verbs create, develop, establish, found, generate, etc.).
To conclude, semantic frames can provide valuable insights into the semantic categorizations of support tickets used by an organization over two or more time periods. They can also help in finding similarities and differences between different pieces of text, even when data availability is limited. As a result, terminology consistency can be ensured and thereby both internal and external communication improved.