Purpose
The Hebrew AI Production Quality Index is intended to provide a repeatable way to evaluate whether Hebrew AI systems are reliable for real workflows rather than merely fluent in isolated examples.
Proposed dimensions
| Dimension | What it measures |
|---|---|
| Hebrew language quality | Morphology, agreement, register, terminology, and naturalness |
| Task correctness | Whether the requested operation is completed correctly |
| Factual reliability | Unsupported claims, reference fidelity, and uncertainty handling |
| Instruction compliance | Following explicit requirements and output formats |
| Safety and refusal | Correct handling of risky and benign requests |
| Formatting and RTL | Mixed-direction text, tables, lists, numbers, and structured outputs |
| Consistency | Behavior across equivalent prompt variants and repeated runs |
| Operational quality | Latency, malformed outputs, missing fields, and integration failures where measurable |
Publication standard
- Disclose test-set scope and limitations.
- Separate model version and configuration.
- Publish scoring rules and severity definitions.
- Use native Hebrew review.
- Report confidence, disagreement, and exclusions.
- Do not present a single composite score without component results.
Next stage
A future public release may include a benchmark dataset, rubric, and report when the test set is sufficiently representative and publication rights are clear. Until then, this page serves as a transparent methodology framework.