Skip to content

Navigation Menu

Sign in
Appearance settings

Search code, repositories, users, issues, pull requests...

Provide feedback

We read every piece of feedback, and take your input very seriously.

Saved searches

Use saved searches to filter your results more quickly

Appearance settings
View jknafou's full-sized avatar

Highlights

  • Pro

Block or report jknafou

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
jknafou/README.md

👋 Hi, I'm Julien Knafou, Ph.D.

Data Scientist | NLP Researcher | Multilingual AI Enthusiast
Geneva, Switzerland


🚀 About Me

I'm a postdoctoral researcher specializing in large language models (LLMs), cross-lingual/domain transfer learning, and machine translation. My work focuses on making scientific and technical knowledge accessible across languages and domains, with a strong emphasis on open-source contributions and reproducible research.


🔬 Research Interests

  • Large Language Models (LLMs)
  • Cross-Lingual & Domain Transfer Learning
  • Machine Translation (MT)
  • Information Retrieval (IR) & Retrieval-Augmented Generation (RAG)
  • Data Augmentation for NLP

🛠️ Skills & Tools

  • Languages: Python, R, LaTeX
  • Frameworks/Libraries: Transformers, PyTorch, TensorFlow, HuggingFace, Ray Tune, FAIRSEQ, FastAPI
  • Infrastructure: Google Cloud, OpenStack, SLURM (HPC)
  • Deployment: Docker, FastAPI
  • Languages Spoken: English, French, Spanish, Portuguese, German (B2)

📚 Selected Projects

  • TransBERT: Leading the development of a state-of-the-art LLM leveraging synthetically translated data. Achieves SOTA performance with open-source code and a forthcoming EMNLP2025 paper.
  • Large Scale Corpus Translation: Designed a scalable package for translating massive corpora (e.g., 22M PubMed abstracts EN→FR), to be released with TransBERT.
  • RAG TREC: Developing a Retrieval-Augmented Generation pipeline for the upcoming TREC conference.

📝 Publications


📫 Let's Connect


"Bridging language barriers through AI and open science."

Pinned Loading

  1. TransCorpus TransCorpus Public

    TransCorpus is a scalable toolkit for large-scale, parallel translation and preprocessing of text corpora, built for language model pretraining and research.

    Python

Morty Proxy This is a proxified and sanitized view of the page, visit original site.