Sections Concept Analyzer FAQ
Manual informationConcept Analyzer FAQ
Concept Analyzer FAQ#
Viewpoint Concept Analyzer (CA) is an Advanced Tool accessible in Viewpoint Review. Its purpose is to identify the concepts present in a set of documents.
- What is a concept?
A concept is simply the assumed topic of a document.
- Why is Conceptual Analysis useful?
It is useful to analyze concepts in order to organize documents by case relevancy. By looking at concept labels you can quickly judge how much weight should be put into reviewing certain concept clusters.
- How does the CA engine label each concept cluster?
To label each cluster, the concept engine chooses an individual word or phrase (sequences of words) from the inputted documents' content. Each document assigned to a concept cluster must contain all words from the cluster's label. Documents containing only some of a cluster label's words will not be assigned to the cluster. As much as possible, the engine will ensure that documents do not overlap too much across concept clusters.
- How does the engine pick which sequence of words to use as a cluster's label?
In general, for each sequence of words found in the input documents, the concept engine computes an aggregated score based on many factors such as the number of occurrences of the phrase, its length, grammatical structure and many others. Then, the highest-scoring phrases become cluster labels. It is important to note that in regards to emails, the concept engine puts more weight in words/phrases appearing in the subject field than in the body of an email.
- How many times must a specific sequence of words appear in a document for it to be considered a concept label?
There is no one threshold that applies to all labels. When clustering large collections of documents the concept engine applies a threshold on the minimum number of occurrences on words and phrases. Words/phrases below these limits are ignored during concept building.
- Can documents be assigned to multiple concept clusters?
A concept cluster contains all documents that contain all of the cluster label's words. This means that if for some pair of cluster labels there are documents that contain both labels, the document will be assigned to two concepts. Obviously, this can also happen for triples but usually does not exceed much more.
- How can I determine which concepts are of higher value?
You could assume that the best concepts are those with the largest number of documents, meaning something along the lines of "dominating concepts". Or conversely, you could assume the best concept is the one that corresponds to the cluster with the smallest number of documents, in which case you would favor the most "specific" topic. Resultantly, concept value cannot be assigned by the engine. It is an abstract value determined by the facts of a case.
Note: The back end concept clustering engine is Carrot Search developed by Lingo3G.
