You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Why is this evaluation interesting?
ArmBench-LLM adds standardized Armenian-language evaluation across 24 tasks, including classification, QA, reasoning, summarization, translation, NER, and language correction. Armenian currently has very limited coverage in common LLM evaluation frameworks.
How used is it in the community?
The benchmark was released recently, so adoption is still growing. Its dataset, results, and evaluation code are publicly available, and the proposed LightEval integration is implemented in PR Add Armenian language evaluation suite (ArmBench-LLM) #1289.
Evaluation short description
Why is this evaluation interesting?
ArmBench-LLM adds standardized Armenian-language evaluation across 24 tasks, including classification, QA, reasoning, summarization, translation, NER, and language correction. Armenian currently has very limited coverage in common LLM evaluation frameworks.
How used is it in the community?
The benchmark was released recently, so adoption is still growing. Its dataset, results, and evaluation code are publicly available, and the proposed LightEval integration is implemented in PR Add Armenian language evaluation suite (ArmBench-LLM) #1289.
Evaluation metadata
Related PR: #1289