Viewpoint
Sign in
Sections Assisted Review Performance Metrics
Manual informationAssisted Review Performance MetricsApplication version 7.0 December 16, 2021 · Document version 7.0 (August 2021)
Source title

Viewpoint™ Assisted Review Interpreting & Using Performance Metrics

Application version

Application Version: 7.0 December 16, 2021

Document version

Document Version: 7.0 (August 2021).

Rights and trademarks

© 2021 Conduent, Inc. All rights reserved. Conduent and Conduent Agile Star are trademarks of Conduent, Inc. and/or its subsidiaries in the United States and/or other countries. Other company trademarks are also acknowledged.

Latest revision

1.1 | August 10, 2021 | Rebranded adhering to the latest Conduent brand central documentation standards/guidelines. | Technical Writer

Document conventions

Convention

Explanation

Bold

For file names, commands, fields, menus, options, and window names.

Courier New

Commands as you should type them.

Lucida Console

Example output generated by the system.

Italics

For configuration variables, including variable portions of file names and URLs. Also indicates a document name.

Note / Blue Callout

The Blue Callout text indicates information that is of special interest or importance, an idea that could be useful or additional information about a product or a feature.

The Caution icon along with the text indicates actions that can lead to problems in system operation or configuration settings if the instructions are not followed properly.

Revision history

This section tracks the initial creation of the document after each major version thereafter.

Ver:

Date

Description

Reviewed / Approved By

1.0

Feb 28, 2020

Initial Version

Team

1.1

August 10, 2021

Rebranded adhering to the latest Conduent brand central documentation standards/guidelines.

Technical Writer

Assisted Review Metrics#

Introduction#

The Performance Metrics grid is a panel available in Viewpoint’s Assisted Review (“VAR”) module. This provides critical information to a user regarding the effectiveness of their assisted review processes. This enables the user to make informed decisions regarding how best to utilize VAR’s predictions; whether it be for the purpose of review prioritization, culling of non-relevant documents, review quality control, etc.

While it is possible to utilize VAR without using Performance Metrics, understanding Performance Metrics information is critical to being able to make defensible decisions about a Viewpoint Assisted Review project.

For example, the Performance Metrics grid can be used to estimate the following information about an unreviewed Target View of documents:

  • The estimated number of relevant documents contained in the Target View
  • Estimated Recall scores associated with specific Probability Score ranges
  • Estimated Precision scores associated with specific Probability Score ranges
  • The estimated number of documents needed to be reviewed to achieve specific Recall or Precision scores
  • The number of relevant documents that are estimated to be ‘left behind’ based on specific Probability Score thresholds, when using VAR to cull not relevant documents from the review corpus

Using This Document#

This document assumes that you are familiar with the Viewpoint Assisted Review module, workflow and terminology. You are advised to review the Viewpoint Assisted Review Manual before using this document.

Using Performance Metrics#

The Performance Metrics grid is used to analyze the effectiveness of a predictive model (based on the Seed Set) by comparing the performance of the model against a reviewed, representative sample of the Target View. Since the analysis is conducted against a representative sample of the Target View, the results of the analysis may be used to generally understand how a Viewpoint Model will likely perform against an unreviewed Target View.

Performance Metrics Requirements#

To utilize the Performance Metrics effectively, the following is required:

  • A fully reviewed, representative sample of documents taken from the Target View. This is typically the random sample of 1000-2000 documents generated at the outset of the Viewpoint Assisted Review process.
  • A predictive Model generated from a Seed View.
  • Probabilities must be calculated against the Sample View using the selected Model.

Note: It is recommended that documents only be tagged with the Positive or Negative tags used during the review of the Sample and Seed Views. Production Tags other than those used as the Positive or Negative tags will not be utilized for the Performance Metrics analysis and may lead to suboptimal results.

Performance Metrics Workflow#

While it is possible to view the Performance Metrics grid on any View in the Assisted Review model that has had Calculate Probabilities run against it, there is a specific workflow that is required in order to get valid Performance Metrics results.

  1. Create Sample from Target View - Typically, the Sample View used for Performance Metrics analysis is the same Sample View generated at the outset of the Viewpoint Assisted Review workflow. It is possible to use other samples for Performance Metrics, but generally Conduent Legal and Compliance Solutions recommends that it has the following attributes:
    • Composed of at least 1,000 documents (usually 1,000 – 2,000)
    • Randomly selected from the Target View

Note: An easy way to generate a random sample from the Target is to use the Actions > Sample Splitter – Number of Documents menu option in the Viewpoint Assisted Review module. Further information about this action is contained in the Viewpoint Assisted Review Manual.

  1. Review Sample View - Since Performance Metrics compares Model-predicted document Probability Scores against reviewer-coded documents, the Sample View must be fully reviewed before the analysis is generated. The review should be conducted using the same Positive and Negative tags used to review the Seed View and generate the Model. Accuracy of the coding is critical as small mistakes in the review of the Sample will be magnified when used to make generalizations about the unreviewed Target View.
  2. Calculate Probabilities on Sample View - Once the review of the Sample View is completed, select the Sample View as the Target View in the Assisted Review module and use the Calculate Probabilities button to apply the Model to the Sample View and generate Probability scores for the sample documents.
  3. Review the Performance Metrics of the Sample - This is described in the “Using Performance Metrics” section.
  4. Generalize Results back to Target View - Since the sample is a randomly drawn, representative sample of the Target View, the Performance Metrics results of the sample will generally represent how the Model would perform against the entire Target View.

For example, if the Performance Metrics against the reviewed Sample indicates that there is 70% recall corresponding to a Probability Score threshold of 50%, then it can be estimated that a 50% score threshold within the Target View will also have an 70% recall.

Important Note: When using the Performance Metrics of a Sample to make generalizations or estimates about the Target View it is important to realize that these measurements will have some associated Sampling Error which is represented by a Margin of Error and Confidence Level. The size of this Sampling Error is a function of the size of the Sample vs the size of the Target View and the underlying ratio of Positive to Negative documents in the Target View. For this reason, it must be understood that the Performance Metrics results are estimates only and do not include the Margin of Error or Confidence Level.

Interpreting Performance Metrics#

Performance Metrics are a powerful feature that enables users to utilize a relatively small sample of reviewed documents in order to make defensible decisions about the disposition of a much larger set of unreviewed documents.

Looking at some typical Performance Metrics taken from a sample of 1,300 documents drawn from a Target View of 1,000,000 documents to see what information is presented:

Probability Range

Documents in Slice

Positive Documents in Slice

% Positive Per Slice

% of Total Positive

Cumulative % of Sample

Cumulative % Positive

Cumulative % of Total Positive

100%

120

113

94.17%

16.19%

9.49%

94.17%

16.19%

100%

61

46

75.41%

6.59%

14.31%

87.85%

22.78%

99.93 - 100%

60

54

90.00%

7.74%

19.05%

88.38%

30.52%

98.2 - 99.91%

60

57

95.00%

8.17%

23.79%

89.70%

38.68%

96.15 - 98.14%

60

54

90.00%

7.74%

28.54%

89.75%

46.42%

93.58 - 96.13%

60

54

90.00%

7.74%

33.28%

89.79%

54.15%

90.49 - 93.56%

60

51

85.00%

7.31%

38.02%

89.19%

61.46%

87.59 - 90.49%

60

44

73.33%

6.30%

42.77%

87.43%

67.77%

83.43 - 87.57%

60

48

80.00%

6.88%

47.51%

86.69%

74.64%

75.92 - 83.42%

60

36

60.00%

5.16%

52.25%

84.27%

79.80%

69.12 - 75.86%

60

36

60.00%

5.16%

57.00%

82.25%

84.96%

60.92 - 69.11%

60

28

46.67%

4.01%

61.74%

79.51%

89.97%

51.81 - 60.82%

60

21

35.00%

3.01%

66.48%

76.34%

91.98%

41.92 - 51.79%

60

20

33.33%

2.87%

71.23%

73.47%

94.84%

29.58 - 41.45%

60

15

25.00%

2.15%

75.97%

70.45%

96.99%

17.73 - 29.56%

60

9

15.00%

1.29%

80.71%

67.19%

98.28%

6.98 - 17.47%

61

4

6.56%

0.57%

85.53%

63.77%

98.85%

1.12 - 6.25%

61

1

1.64%

0.14%

90.36%

60.45%

99.00%

0 - 1.12%

61

6

9.84%

0.86%

95.18%

57.89%

99.86%

0 - 0%

61

1

1.64%

0.14%

100.00%

55.18%

100.00%

 

1,265

698

 

Sample Performance Metrics

Report Format#

The Performance Metrics data is presented in what is often called a “Lift” table. This table takes all the documents in the selected Target View (in this case our random Sample View) and ranks them according to their Probability Score; from highest to lowest. The table then divides the documents into approximately 20 slices of the ranked data, each representing roughly 5% of the total data in the selected View. Each slice is represented as a row in the table along with the corresponding Probability Score range associated with that slice.

Note: You may note that the total of the documents in the Performance Metrics report may not equal the total of the documents in the selected View. The Performance Metrics information will not include documents that did not receive predictions because they did not contain text, contained too few/too many tokens, etc.

Note: Some slices may represent significantly more than 5% in circumstances where documents possessing one single score may represent more than 5% of the data. This will typically happen only with scores around 100% or 0%. Because of this, there will sometimes be less than 20 rows present in the table.

Report Column Descriptions#

  • Probability Range: The Probability Score range associated with the row (or “slice”) of the documents in the selected View. The Probability Score range for each row will decrease going down the table.
  • Documents in Slice: The number of documents contained in the slice. This will generally be approximately 5% of the total number of documents in the selected View, excepting those situations described in Footnote 4 below.
  • % Positive per Slice: The number of reviewer-coded documents in the slice that were coded with the Positive Production Tag.
  • % of Total Positive: The number of reviewer coded Positive documents in the slice as a percentage of all the total Positive documents in the selected View.

The last three rows of the table – Cumulative % of Sample, Cumulative % Positive, and Cumulative % of Total Positive are generally regarded as the three most important columns of the table and operate a bit different than the other columns since they are presenting cumulative data. What this means is that the values in these cells are cumulative of the selected row and all the rows above that point in the table. Let’s take a look at the last three columns of our table on the next page:

To help explain these columns, let’s assume that, for an example review protocol, it has been agreed that the protocol will seek to achieve an 80% recall of the estimated volume of relevant (Positive) documents in the review corpus (the Target View).

Probability Range

Cumulative % of Sample

Cumulative % of Positive

Cumulative % of Total Positive

100%

9.49%

94.17%

16.19%

100%

14.31%

87.85%

22.78%

99.93 - 100%

19.05%

88.38%

30.52%

98.2 - 99.91%

23.79%

89.70%

38.68%

96.15 - 98.14%

28.54%

89.75%

46.42%

93.58 - 96.13%

33.28%

89.79%

54.15%

90.49 - 93.56%

38.02%

89.19%

61.46%

87.59 - 90.49%

42.77%

87.43%

67.77%

83.43 - 87.57%

47.51%

86.69%

74.64%

75.92 - 83.42%

52.25%

84.27%

79.80%

69.12 - 75.86%

57.00%

82.25%

84.96%

60.92 - 69.11%

61.74%

79.51%

88.97%

51.81 - 60.82%

66.48%

76.34%

91.98%

41.92 - 51.79%

71.23%

73.47%

94.84%

29.58 - 41.45%

75.97%

70.45%

96.99%

17.73 - 29.56%

80.71%

67.19%

98.28%

6.98 - 17.47%

85.53%

63.77%

98.85%

1.12 - 6.25%

90.36%

60.45%

99.00%

0 - 1.12%

95.18%

57.89%

99.86%

0 - 0%

100.00%

55.18%

100.00%

Looking at these columns from right to left:

  • Cumulative % of Total Positive: This value indicates that, if all the documents in the selected row were reviewed, plus all the documents in the rows above it, the review process should result in having reviewed this percentage of all the Positive documents in the selected View. This value corresponds to the Recall measurement associated with that row’s score threshold.

In the example protocol, it is agreed that the review process should seek to achieve an 80% Recall, which approximately corresponds to the value of the Cumulative % Positive value of 79.80% contained in the 10th row in the chart on the right. Since it is a cumulative score, this means that all the rows preceding the 10th row would have to be reviewed in order to achieve this Recall score.

In order to achieve an 80% recall for our review, it would be necessary to review all documents possessing a Probability Score range of 75.92 – 100%.

  • Cumulative % Positive: This value indicates that, if all the documents in the selected row were reviewed, plus all the documents in the rows above it, this is the total percentage of documents that would be Positive out of all the documents reviewed. This value corresponds to the Precision measurement associated with that row’s score threshold.

According to the example protocol, if all documents with a Probability Score range of 75.92% and above are reviewed, the reviewed set would have a Precision of approximately 84% (meaning 84% of the reviewed documents would be Positive.)

  • Cumulative % of Sample: This value indicates that, if all the documents in the selected row were reviewed, plus all the documents in the rows above it, the review protocol would end up reviewing this percentage of the total document corpus.

Referring back to the example protocol again, a Recall threshold of 80% has a Cumulative % of Sample value of 52.25%. This means that it would be necessary to review 52.25% of the example document corpus in order to achieve the desired Recall level.

Applying Sample Performance Metrics to the Review Corpus#

Knowing that a reviewed random sample of a document corpus is generally representative of that corpus, this means that Performance Metrics from a reviewed sample may be used to make informed decisions about the unreviewed document corpus (the original Target View). However, it is important to understand that when applying sample metrics to the Target View, one must keep in mind that there is some amount of Sampling Error associated with that measurement.

The following chart gives an example of how Sampling Error operates when comparing Sample metrics back to the original Target View using the Performance Metrics from our previous tables:

  • Size of Target View: 1,000,000 documents
  • Size of Reviewed Sample: 1,265 documents
  • Positive Documents in Sample: 698 (~55% of Sample)
  • Sampling Error: 95% Confidence Level with ~2.75% Margin of Error
  • Desired Recall: 80%

Sample Performance Metrics Applied to Target View

Metric

As Measured by Sample (rounded)

As Applied to Target View

Recall (%)

80%

80 ± 2.75%

Precision (%)

84%

84 ± 2.75%

Corpus to Review (%)

52%

52%

Corpus to Review (#)

658

520,000

Note: Viewpoint does not directly calculate the Sampling Error associated with the Sample View. Sampling Error is a function of the size of the Sample, the size of the Target View, and the underlying relevance rate of the Target View (as measured by the Sample.) A good sampling calculator is available online at http://www.surveysystem.com/sscalc.htm. Conduent Legal and Compliance Solutions does not provide any guarantees regarding the accuracy of the results when using this website.