An open, transparent framework for evaluating AI model trustworthiness.
K-BENCHMARK is ALPAR AI's open methodology for scoring AI models across safety, truthfulness, fairness, privacy, robustness, and transparency dimensions. Scores are computed from verified incident reports, cross-audit engine results, and domain expert evaluations using Wilson-score confidence intervals.
Real-world AI incident evaluation grounded in the ALPAR incident registry, combined with adversarial jailbreaks.
TruthfulQA-based factuality analysis paired with curated Turkish-context factual claims.
Evaluation of Turkish language proficiency using translated MMLU-TR and bespoke local regulatory test sets.
Assessment of model decision-making logic against Art. 73 regulatory scenarios based on EU AI Act taxonomy.
Standard mathematical problem-solving capability checked against GSM8K and MATH benchmarks.
Verification of structural instruction-following accuracy using IFEval-like constraints.
Adversarial jailbreak protection and prompt-injection defense rating under human-approved simulation.
Needle-in-a-haystack retrieval evaluation for context lengths up to 32k tokens.
Each category is scored using a Wilson-score confidence interval, which accounts for sample size and provides a statistically rigorous lower bound. This prevents small-sample noise from inflating ratings.
Wilson score = (p + z²/2n - z√(p(1-p)/n + z²/4n²)) / (1 + z²/n)
Wilson-score is a standard method in binomial proportion estimation, commonly used in platforms like Reddit and IMDb for robust rating systems.
Incident Submission
PII Guard
Cross-Audit Engine
Wilson Scoring
Expert Review
Scores are derived from: (1) verified incident reports submitted via the ALPAR platform, (2) automated cross-audit engine evaluations using established benchmarks (MMLU, GSM8K, IFEval, BBH), (3) domain expert assessments from the L3 advisor network, and (4) publicly available model documentation and technical reports.
K-BENCHMARK scores are informational and provided 'as-is'. They do not constitute a compliance certification or legal opinion. ALPAR AI is 'AI Act ready/aligned', not compliant. See Terms of Service for full disclaimer.