Skip to content

Navigation Menu

Sign in
Appearance settings

Search code, repositories, users, issues, pull requests...

Provide feedback

We read every piece of feedback, and take your input very seriously.

Saved searches

Use saved searches to filter your results more quickly

Appearance settings
#

language-model-evaluation

Here are 10 public repositories matching this topic...

Language: All
Filter by language

Layer-wise Semantic Dynamics (LSD) is a model-agnostic framework for hallucination detection in Large Language Models (LLMs). It analyzes the geometric evolution of hidden-state semantics across transformer layers, using contrastive alignment between model activations and ground-truth embeddings to detect factual drift and semantic inconsistency.

  • Updated Mar 12, 2026
  • Python

Probed gender-science bias in transformer language models by implementing WEAT on English and Urdu embeddings. Conducted comparative evaluation across BERT, RoBERTa, DistilBERT, and XLM-RoBERTa, revealing model and language-dependent bias patterns and providing reproducible benchmarks.

  • Updated Feb 21, 2026
  • Jupyter Notebook

Evaluation framework for measuring LLM reasoning and translation performance across low-resource languages using glossary-guided prompting and semantic similarity metrics.

  • Updated Jun 19, 2026
  • Python

Improve this page

Add a description, image, and links to the language-model-evaluation topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the language-model-evaluation topic, visit your repo's landing page and select "manage topics."

Learn more

Morty Proxy This is a proxified and sanitized view of the page, visit original site.