Research framework

Hebrew AI Production Quality Index

A transparent evaluation framework for production Hebrew AI systems. The current public page defines the proposed dimensions and publication standard; it does not claim that a market-wide score has already been completed.

Purpose

The Hebrew AI Production Quality Index is intended to provide a repeatable way to evaluate whether Hebrew AI systems are reliable for real workflows rather than merely fluent in isolated examples.

Proposed dimensions

DimensionWhat it measures
Hebrew language qualityMorphology, agreement, register, terminology, and naturalness
Task correctnessWhether the requested operation is completed correctly
Factual reliabilityUnsupported claims, reference fidelity, and uncertainty handling
Instruction complianceFollowing explicit requirements and output formats
Safety and refusalCorrect handling of risky and benign requests
Formatting and RTLMixed-direction text, tables, lists, numbers, and structured outputs
ConsistencyBehavior across equivalent prompt variants and repeated runs
Operational qualityLatency, malformed outputs, missing fields, and integration failures where measurable

Publication standard

  • Disclose test-set scope and limitations.
  • Separate model version and configuration.
  • Publish scoring rules and severity definitions.
  • Use native Hebrew review.
  • Report confidence, disagreement, and exclusions.
  • Do not present a single composite score without component results.

Next stage

A future public release may include a benchmark dataset, rubric, and report when the test set is sufficiently representative and publication rights are clear. Until then, this page serves as a transparent methodology framework.