# ReliToolbox Validation Roadmap

This document defines the short-term validation framework for ReliToolbox. Its purpose is to make every reliability calculation traceable, testable, and clearly bounded.

ReliToolbox is intended for engineering screening, workflow support, and educational reliability analysis. It is not a certified replacement for enterprise reliability, safety, RBI, or SIL verification software unless a module has been independently verified and formally qualified for that use.

The first benchmark validation cases are maintained in [benchmarks/VALIDATION_CASES.md](benchmarks/VALIDATION_CASES.md). The initial set covers Weibull, RAM, RBD, and FTA.

---

## 1. Validation Philosophy

Every module should have:

1. A clearly stated engineering purpose.
2. A documented mathematical model.
3. A reference calculation or benchmark case.
4. Defined input units and valid ranges.
5. Expected output values with numerical tolerance.
6. Known limitations and invalid use cases.
7. A confidence level.

The goal is not to claim enterprise certification. The goal is to make the tool reliable enough for early engineering judgment, internal discussion, teaching, and preliminary screening.

---

## 2. Confidence Levels

| Level | Name | Meaning | Typical Use |
|---|---|---|---|
| Level 1 | Educational / Screening | Formula is reasonable and useful for learning or early screening, but not yet fully benchmarked. | Training, quick checks, concept comparison |
| Level 2 | Engineering Support | Core calculation has been benchmarked against published examples or independent hand calculations. Boundary conditions are documented. | Internal engineering review, pre-study, option comparison |
| Level 3 | Design / Audit Ready | Calculation has a complete validation pack, peer review, regression tests, version control, and traceable references. | Formal design support only after independent review |

Default status: unless otherwise stated, modules should be treated as Level 1 or Level 2 tools, not as certified design or audit tools.

---

## 3. Module Validation Matrix

| Module | Primary Method | Current Target Level | Required Validation Cases | Key Risks to Check |
|---|---|---:|---|---|
| Weibull | MLE fitting, censoring, B-Life | Level 2 | Complete failures, right-censored data, small sample, B10/B50 | Wrong censoring likelihood, eta/beta convention, CI interpretation |
| Weibull+ | Interval censoring, confidence intervals | Level 2 | Interval-censored example, mixed censoring, CI benchmark | Incorrect likelihood, unstable optimization, misleading confidence band |
| AMSAA | Crow-AMSAA / NHPP reliability growth | Level 2 | Known beta/lambda example, Laplace trend case | Calendar time vs cumulative test time confusion |
| ALT | Arrhenius acceleration | Level 1-2 | Single stress extrapolation, multiple stress levels | Celsius/Kelvin error, over-extrapolation |
| Degradation | Degradation trajectory and threshold crossing | Level 1-2 | Linear degradation, nonlinear trend, noisy data | Extrapolation beyond data range, wrong failure threshold direction |
| RBD | Reliability block diagram | Level 2 | Series, parallel, k-out-of-n, mixed system | Incorrect independence assumption, path/cut set logic errors |
| RBD/MCS | Minimal cut set analysis | Level 2 | Simple FTA/RBD converted cases | Duplicate/minimal set reduction errors |
| RAM | Event-driven Monte Carlo | Level 2 | 1oo1 repairable system, parallel repairable system, capacity case | TTF/TTR scheduling, warm/cold standby, capacity propagation |
| FTA | Fault tree, MOCUS, importance | Level 2 | AND, OR, mixed gates, repeated events | Non-minimal cut sets, repeated basic events, importance formula errors |
| Allocation | Equal, weighted, AGREE allocation | Level 1-2 | Simple system target allocation | Mission time and weighting interpretation |
| FMEA/RCM | RPN/AP, action closure, RCM task suggestion | Level 1-2 | RPN ranking, AP category, task mapping | RPN misuse, action priority overclaim, missing detectability limits |
| HAZOP | Guideword worksheet and risk matrix | Level 1 | Node/parameter/deviation examples | Treating worksheet as complete HAZOP without team review |
| SIL | IEC 61508/61511 PFD/PFH approximations | Level 2 | 1oo1, 1oo2, proof test interval, beta factor | Dangerous failure assumptions, low/high demand confusion |
| LOPA | Initiating event frequency, IPL reduction, SIL need | Level 2 | Single initiating event, multiple IPLs | Non-independent IPLs, misuse of enabling conditions |
| Event Tree | Barrier branch frequency | Level 1-2 | Binary branch example, multiple consequence paths | Probability sum errors, barrier dependence |
| Bowtie | Threat-barrier-consequence diagram | Level 1 | Standard bowtie example | Diagram treated as quantitative proof without data |
| RBI | Semi-quantitative PoF x CoF | Level 1-2 | 5x5 risk matrix, inspection interval example | Misstating API 581 compliance, corrosion rate uncertainty |
| Predict | MIL-HDBK-217F-style parts count | Level 1-2 | Basic electronic component count example | Data vintage, environment factor misuse |
| Stress-Strength | Stress-strength interference | Level 2 | Normal-normal case, lognormal case, Monte Carlo check | Distribution mismatch, tail probability numerical error |
| Demo Test | MTBF demonstration, chi-square method | Level 2 | Zero-failure and failure-allowed cases | One-sided/two-sided confidence confusion |
| LCC | NPV, EAC, alternative comparison | Level 2 | Discounted cash flow benchmark | Discount timing, nominal vs real cost |
| Spares | Poisson demand / service level | Level 2 | Low demand spare, target fill probability | Demand independence, repairable vs consumable confusion |
| PM Opt | Age replacement cost-rate model | Level 2 | Weibull age replacement example | Wrong renewal cycle cost, beta <= 1 interpretation |
| Equipment | Equipment registry, FRACAS, inspection plan | Level 1-2 | Import/export, failure linkage, dashboard consistency | Data loss, duplicate tags, inconsistent foreign keys |
| Health | Data health and workflow gaps | Level 1-2 | Orphan failure, overdue RBI, missing Weibull | False positives/negatives in gap logic |
| CSV Preflight | CSV validation before import | Level 2 | Missing columns, duplicate tags, invalid numbers | Silent import errors, type coercion |

---

## 4. Minimum Benchmark Record Format

Each benchmark case should be documented using the following structure:

```text
Benchmark ID:
Module:
Purpose:
Reference:
Inputs:
Expected Outputs:
Tolerance:
Assumptions:
Known Limitations:
Validation Status:
```

Example:

```text
Benchmark ID: WEIBULL-COMPLETE-001
Module: Weibull
Purpose: Verify two-parameter Weibull MLE for complete failure data.
Reference: Independent hand calculation or published textbook example.
Inputs: Failure times, all exact failures, no censoring.
Expected Outputs: beta, eta, MTTF, B10.
Tolerance: beta and eta within 0.5%; B-Life within 1%.
Assumptions: Independent and identically distributed failure times.
Known Limitations: Does not validate mixed censoring or confidence intervals.
Validation Status: Planned.
```

---

## 5. Unit and Terminology Rules

Reliability calculations are highly sensitive to unit and definition errors. Each module should explicitly show units for:

| Quantity | Preferred Display |
|---|---|
| Time to failure | h, day, year; never mixed silently |
| Failure rate | 1/h or failures/year |
| Repair time | h |
| Availability | dimensionless or % |
| Probability | dimensionless, 0 to 1 |
| Frequency | events/year |
| Cost | user-selected currency with no implicit conversion |
| Temperature for Arrhenius | Kelvin internally; Celsius only for input display |

Terminology should be consistent:

- MTBF is for repairable systems or repeated failure events.
- MTTF is for non-repairable life distribution context.
- Eta is Weibull characteristic life, not MTBF unless a specific relationship is calculated.
- Beta is the Weibull shape parameter.
- B10 is the time at which unreliability reaches 10%, i.e. reliability is 90%.
- PFDavg and PFH should not be mixed.
- LOPA IPLs must be independent, auditable, specific, and effective.

---

## 6. Standard Disclaimer for Modules

Recommended module footer text:

> This module is intended for engineering screening and educational analysis. Results should be independently verified before use in design, safety, inspection, maintenance, or commercial decisions. Confirm units, assumptions, data quality, and applicable standards before relying on the output.

For SIL, RBI, LOPA, and HAZOP modules, use a stronger footer:

> This module does not replace a formal IEC 61511 SIL verification, API 580/581 RBI assessment, LOPA study, or HAZOP team review. It provides structured calculation support and preliminary screening only.

---

## 7. Short-Term Implementation Checklist

Priority actions:

1. Add validation status to each module page.
2. Add one benchmark case per core module.
3. Add unit labels beside all numerical inputs and outputs.
4. Add invalid-input checks for negative time, zero demand rate, probability outside 0-1, and missing equipment links.
5. Add exportable validation report for benchmark results.
6. Keep benchmark expected results under version control.
7. Add a README link to this validation roadmap.

Core module priority:

1. Weibull
2. RAM
3. RBD
4. FTA
5. SIL
6. LOPA
7. RBI
8. PM Opt
9. Spares
10. LCC

---

## 8. Practical Acceptance Criteria

A module may be treated as Level 2 only when:

1. At least one benchmark case exists.
2. The benchmark has expected numerical output.
3. The result can be reproduced after code changes.
4. Units are visible in the UI.
5. Limitations are documented.
6. The module avoids unsupported compliance claims.

A module should remain Level 1 when:

1. It is mainly a worksheet or workflow tool.
2. It has not been benchmarked.
3. It relies on simplified risk matrices.
4. It depends heavily on expert judgment.
5. It implements only a partial version of an industry standard.

---

## 9. First Benchmark Baseline Status

The first benchmark baseline has been added in [benchmarks/VALIDATION_CASES.md](benchmarks/VALIDATION_CASES.md).

| Module | Benchmark Cases Added | Numerical Expected Values Locked | Current Validation Status |
|---|---:|---:|---|
| Weibull | 2 | Partial | Level 1 moving toward Level 2 |
| RAM | 2 | Yes | Level 2 baseline for simple repairable systems |
| RBD | 3 | Yes | Level 2 baseline for deterministic logic |
| FTA | 3 | Yes | Level 2 baseline for basic gate and cut-set logic |

Weibull still requires independent numerical reference values for beta, eta, MTTF, B10/B50/B90, and confidence intervals before it should be fully marked as Level 2.

---

## 10. Engineering Positioning Statement

ReliToolbox should be positioned as:

> A lightweight, browser-based reliability engineering toolbox for screening, education, workflow integration, and preliminary engineering support.

It should not be positioned as:

> A certified replacement for enterprise RAM, RBI, SIL, FMEA, or safety lifecycle software.

This distinction protects users from over-reliance and makes the tool more credible.
