Guidance for Industry and FDA Staff

PDF Printer Version Document issued on: March 23, 2011

For questions regarding this document contact Donna Roscoe at 301-796-6183 ([email protected]), or Marina Kondratovich at 301-796-6036 ([email protected]).

CDRH Logo

U.S. Department of Health and Human Services
Food and Drug Administration
Center for Devices and Radiological Health
Office of In Vitro Diagnostic Device Evaluation and Safety
Division of Immunology and Hematology Devices

Preface

Public Comment

You may submit written comments and suggestions at any time for Agency consideration to the Division of Dockets Management, Food and Drug Administration, 5630 Fishers Lane, rm. 1061, (HFA-305), Rockville, MD, 20852. Submit electronic comments to http://www.regulations.gov. Identify all comments with the docket number listed in the notice of availability that publishes in the Federal Register. Comments may not be acted upon by the Agency until the document is next revised or updated.

Additional Copies

Additional copies are available from the Internet. You may also send an e-mail request to [email protected] to receive an electronic copy of the guidance or send a fax request to 301-827-8149 to receive a hard copy. Please use the document number (1707) to identify the guidance you are requesting.

Table of Contents

  1. INTRODUCTION
  2. BACKGROUND
  3. SCOPE
  4. RISKS TO HEALTH
  5. DEVICE DESCRIPTION 
    1. Background
    2. Quality Systems Regulation (QS Reg)
    3. Intended Use/Indications for Use
    4. Test Rationale
    5. Test Components and Methodology
    6. Test Results
  6. ANALYTICAL PERFORMANCE VALIDATION
    1. Specimen
    2. Repeatability / Reproducibility
    3. Linearity of Individual Analytes
    4. Performance at Low Levels
    5. Interference
    6. Cross-reactivity/non-specific binding
    7. Hook Effect of the Individual Analytes
    8. Carry-Over Contamination
    9. Matrix comparison
    10. Stability
    11. Calibration and Controls
  7. SOFTWARE
  8. CLINICAL PERFORMANCE EVALUATION
    1. Study population/samples
    2. Cut-Off/ Clinical Decision Points
    3. Clinical Reference Standard (“Gold Standard”)
    4. Study Design
    5. Expected Values in Other Benign and Malignant Conditions
    6. Reference Intervals
    7. Relevance of the Individual Analytes Included in the Score
  9. LABELING
  10. REFERENCES

Guidance for Industry and Food and Drug Administration Staff

Class II Special Controls Guidance Document: Ovarian Adnexal Mass Assessment Score Test System

1 You should determine the level of concern prior to the mitigation of hazards. In vitro diagnostic devices of this type are typically considered a moderate level of concern because software flaws could result in false results reported to clinician and patient, which could cause harm to the patient.

You should include the following points, as appropriate, in preparing software documentation for FDA review:

  • Full description of the software design. Your software should not include utilities that are specifically designed to support uses beyond those in your intended use. You should also consider privacy and security issues in your design. Information about some of these issues may be found at the following website regarding the Health Insurance Portability and Accountability Act (HIPAA) http://www.hhs.gov/ocr/privacy/hipaa/understanding/index.html.
  • Hazard analysis based on critical thinking about the device design and the impact of any failure of subsystem components, such as signal detection and analysis, data storage, system communications, and cybersecurity in relationship to incorrect patient reports, instrument failures, and operator safety.
  • Documentation of complete verification and validation (VV) activities for the version of software that will be submitted to demonstrate substantial equivalence. You should also submit information regarding validation of the compatibility of test software with any instrumentation software.
  • If the information you include in the 510(k) is based on a version other than the release version, identify all differences in the 510(k) version and detail how these differences (including any unresolved anomalies) impact the safety and effectiveness of the device.

Below are additional references to help you develop and maintain your device under good software life cycle practices consistent with FDA regulations.

  • General Principles of Software Validation; Final Guidance for Industry and FDA Staff; available on the FDA Web site
  • Guidance for Off-the-Shelf Software Use in Medical Devices; Final; available on the FDA Web site.2
  • 21 CFR 820.30 Subpart C – Design Controls of the Quality System Regulation.
  • ISO 14971-1; Medical devices – Risk management – Part 1: Application of risk analysis.
  • AAMI SW68:2001; Medical device software – Software life cycle processes.

3

You should demonstrate that your test provides additional information for biologically relevant subpopulations (e.g., pre-menopausal, post-menopausal) or provide an acceptable justification for why such a demonstration is not needed.

Consider a general scheme of comparison of test T1 and combination OR of tests T1 and T2. The sensitivity for the “OR” combination is at least as large as the sensitivity for T1 alone. The specificity for the “OR” combination is the same or worse than the specificity for T1 alone. Thus, the combination “OR” has an inherent trade-off between sensitivity and specificity. An increase in combined sensitivity alone does not prove that the combination (OR) of the T1 and T2 tests is effective if the combined specificity is shown to decrease appreciably.

The data of the clinical study should demonstrate that

  • There is a statistically and clinically significant improvement in NPV with the combination OR (pre-surgical assessment and the score test) vs. NPV of the pre-surgical assessment alone; and
  • If there is a loss in PPV with the combination OR of the pre-surgical assessment and the score test vs. PPV of the pre-surgical assessment alone, this loss in PPV should be clinically acceptable.

The logic for these success criteria is developed below. Consider a straight line connecting the point (0,0) through the point corresponding to test T1, on a plot of Sensitivity vs. 1-Specificity (see Figure 1 below). This line denotes performance characteristics for tests that have the same PPV as test T1. A straight line connecting the point (1,1) through the point corresponding to test T1 denotes the performance characteristics of tests that have the same NPV as test T14 In comparing the performance of test T1 with the performance of the OR combination of tests T1 and T2, there are three possible scenarios:

Scenario A: Both predictive values (PPV and NPV) for “OR” combination (T 1 or T 2) are larger than predictive values of T 1 alone (see the green region). Thus, it is easy to draw the conclusion that the combination OR is better than the T 1 alone.

This figure is a graph with the Y-axis labeled as Se (sensitivity) and the X-axis labeled as 1-specificity. A diagonal line (line 1) is shown extending from the origin (x=0, y=0) to the upper right corner of the graph. Another line (line 2) is shown extending from the origin (x=0, y=0) to the point that is at the upper end of the y-axis and halfway across the x-axis (x=0.5, y=1) A third line (line 3) extends from a point that is at (near (y=0.66, x=0) to the upper right corner (x=1, y=1). At the intersection between line 2 and line 3, there is a vertical line upward to y=1 (line 4); and there is a horizontal line across to x=1 (line 5). The area between line 3 and line 4 is shaded green; the area between line 3 and line 5 is shaded orange. The area between the green and orange triangles thus formed is shaded blue, and there is a dot in the center of this area that is labeled T1 OR T2. The area between lines 1, 2, and 5 (unshaded) is labeled T1.

Figure 1

Scenario B: The PPV of combination OR is worse than PPV of T1 but the NPV of combination OR is better than NPV of T1 (see the Blue region). In this region there is a trade-off between the amount by which NPV increased and the amount by which PPV decreased. Success can be concluded if the lowered PPV remains consistent with safe and effective use of the test.

Scenario C: Both PPV and NPV of combination OR are worse than the PPV and NPV of test T1 (see the red region). Thus, it is easy to draw the conclusion that the combination OR is worse than the T1 alone.

We recommend you summarize the results of the clinical study based on pathology results in tables similar to the ones below and provide an assessment of probability of malignancy (along with 95% CI) based on the various outcomes shown in the tables below:

Present the performance of the score test and pre-surgical assessment for the subjects with malignancy by pathology and for subjects with no malignancy by pathology separately.

Malignancy by Pathology

No Malignancy by Pathology

The Table below shows performance characteristics for the test applied to all subjects evaluated by non-GO physicians. For Single Assessment, only the pre-surgical assessment is used, without reference to a score test result. For Dual Assessment (i.e., “OR” combination) the adnexal mass is declared potentially malignant if the pre-surgical clinical assessment, the score test, or both were positive.

  • Provide sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV) along with the 95% confidence intervals for the pre-surgical assessment alone;
  • Provide sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV) for the Score test performance in conjunction with the pre-surgical assessment using the decision rule “OR” along with 95% confidence intervals.
  • Calculate the difference in NPVs and difference in PPVs along with 95% two-sided confidence intervals (the bootstrap technique can be used for calculation of the confidence intervals). Improvement in NPV should be statistically and clinically significant, and, if a loss in PPV is observed, you should justify the clinical acceptability of this loss.

In addition, you should present the observed frequencies of malignancy for different results of the pre-surgical assessment and the Score test results from the patients evaluated by non-GO in the table below along with 95% confidence intervals:

The same information should be presented as likelihood ratios along with their 95% CI, tabulated as illustrated above for frequencies of malignancy. Likelihood ratio (Result) = Pr(Result|Malignancy) / Pr(Result|No Malignancy). Likelihood ratio, unlike predictive value, is independent of the prevalence of disease.

Subgroup Analyses: You should demonstrate statistically and clinically significant improvement in NPV of dual assessment vs. NPV of single assessment for pre-menopausal and post-menopausal patients separately analysis analogous to the one described for the overall population.

Additional Information: You should provide a tabulation of the descriptive statistics for your Score test within patients grouped according to tumor stage or histopathological findings.

Results in Patient Populations Evaluated by Gynecologic Oncologists

Ovarian adnexal mass assessment score test system is intended for women with pelvic masses who will be having surgery. The test is indicated as an aid in making referral decisions. Your clinical study should avoid possible bias in results from evaluating patients who may have been selectively enrolled at non-GO sites. (For example, some non GO physicians may automatically refer patients to a GO for various reasons regardless of the pre-surgical evaluation. The enrollment in your clinical trial at such a site would lead to a potentially biased representation of patients.) You may opt to provide additional data in GO-evaluated patients to demonstrate a positive bias is not occurring in your test performance. This data is reviewed by FDA with the expectation that the performance is not diminished in the GO-evaluated group.

E. Expected Values in Other Benign and Malignant Conditions

The target population may have a wide variety of conditions unrelated to cancer but present at the time an ovarian mass has been identified. These other conditions (for which in some cases the actual measurements of the analytes [e.g., immunoassays that are components of the test] are indicated) could dramatically affect the Score test result and confound its interpretation. You should demonstrate the results of your test score in patients with the disease conditions indicated by the individual analyte assays, as well as benign and malignant conditions that may be occurring concurrently. Examples of these conditions are: endometriosis, pelvic inflammatory disease, diabetes, anemia, autoimmune diseases such as Crohn’s, SLE and rheumatoid arthritis, cardiac disease, hepatitis, kidney diseases and malnutrition, and various cancers such as cervical cancer, lung cancer, breast cancer, and colorectal cancer.

F. Reference Intervals

Reference values in apparently healthy women may be provided, though such women are not part of the intended use population for the Score test. For each analyte and for the Score test result, any reference values should include women that span the age range of your test and should evaluate a minimum of 120 premenopausal women and 120 postmenopausal women unless you are able to demonstrate that there are not any differences between the two populations. You should include other ethnicities if possible (Latino and Asian) in addition to Caucasian and African American. You should provide the score for each woman and investigate the relationship of the score vs. age.

G. Relevance of the Individual Analytes Included in the Score

You should justify inclusion of each individual analyte in the use of the Score test. One option is to demonstrate that the individual analytes included in the calculation of the score are informative for ovarian malignancy using the data from the clinical study. For this, perform ROC analyses: for each individual analyte, present an ROC curve of the individual analyte and calculate the areas under ROC curve of the individual analyte along with confidence interval (multiplicity issue should be properly addressed). In addition, for each individual analyte, present an ROC curve of the individual analyte and the ROC curve of the score on the same graph. If the data of the clinical study did not demonstrate that some individual analytes are informative for ovarian malignancy, you should justify why these analytes were included in the calculation of the Score test.

1 http://www.fda.gov/downloads/MedicalDevices/DeviceRegulationandGuidance/ GuidanceDocuments/ucm073779.pdf

2 http://www.fda.gov/downloads/MedicalDevices/DeviceRegulationandGuidance/ GuidanceDocuments/ucm073779.pdf

3 Biggerstaff, B.J. Comparing diagnostic tests: a simple graphic using likelihood ratios. Statistics in Medicine 2000, 19: 649-663.

4 PPV depends on a positive likelihood ratio (PLR) and prevalence of malignancy, and NPV depends on negative likelihood ratio (NLR) and prevalence. For a comparison of two tests within the same population, comparison of PPV and NPV is equivalent to comparison of PLR and NLR.