Orazio Oztas

MSc student in Computational Linguistics at the University of Turin, working where language meets AI.

My work sits between NLP, corpus linguistics, and how technical communities actually use words. I have a humanities background, a BA in Film & Media Studies (110/110 cum laude) whose thesis on AI in science-fiction cinema is what pulled me toward language models in the first place.

The slightly unusual part: I use AI coding tools (Claude Code, Codex, Antigravity) hard, every day, as a real workflow, and I'm learning the ML underneath from scratch, in public.

What I'm focused on

An MSc thesis on LLM evaluation. The question I keep coming back to: can we trust an LLM-as-a-judge when the judge is small or local, and humans themselves disagree on the answer? It connects my annotation background (inter-annotator agreement) to a problem labs actually have. Plain Python, local models, built in the open.

Selected work

dark-current-mlx
Local, Apple Silicon (MLX) test for hidden position bias in LLM-as-a-judge models.
personal-os
A sanitized, public skeleton of the operating layer I run on top of Claude Code: context architecture, skills, hooks, model-routing.
doc-qa-rag
RAG-based document Q&A API using FastAPI, ChromaDB, and Sentence Transformers.
llm-fine-tuning
LLM fine-tuning pipeline using LoRA and QLoRA with Hugging Face Transformers.
paperradar
Academic paper monitoring tool with AI summarization.
message-to-action-it
Convert natural-language text into actionable items with ICS calendar export.

Also

A corpus-linguistics paper on how the word agent shifted meaning between two ML eras (arXiv abstracts, 2017–2018 vs 2024), a bibliometric case study of the AI and neural-networks field, and reading ML papers one at a time, posting the notes as I go.

Elsewhere