5036 25 8

Albert Villanova del Moral

albertvillanova

https://albertvillanova.github.io/

AI & ML interests

ML Engineer @ Hugging Face: Evaluations (Science)

Recent Activity

updated a Space about 12 hours ago

demo-leaderboard-backend/leaderboard

New activity 1 day ago

demo-leaderboard-backend/leaderboard:Remove user_name and model_path from submit.py

updated a dataset 1 day ago

albertvillanova/tmp-file-upload

View all activity

Organizations

albertvillanova's activity

updated a Space about 12 hours ago

Running on CPU Upgrade

🥇

Demo Leaderboard

New activity in demo-leaderboard-backend/leaderboard 1 day ago

Remove user_name and model_path from submit.py

#16 opened 1 day ago by

albertvillanova

updated a dataset 1 day ago

albertvillanova/tmp-file-upload

Viewer • Updated 1 day ago • 2

New activity in hf-doc-build/doc-build 2 days ago

Create lighteval/_versions.yml

#28 opened 3 months ago by

SaylorTwift

updated a Space 8 days ago

Running

🏆

Open LLM Leaderboard Model Comparator

Compare Open LLM Leaderboard results

posted an update 8 days ago

Post

1137

🚨 How green is your model? 🌱 Introducing a new feature in the Comparator tool: Environmental Impact for responsible #LLM research!
👉 open-llm-leaderboard/comparator
Now, you can not only compare models by performance, but also by their environmental footprint!

🌍 The Comparator calculates CO₂ emissions during evaluation and shows key model characteristics: evaluation score, number of parameters, architecture, precision, type... 🛠️
Make informed decisions about your model's impact on the planet and join the movement towards greener AI!

updated a dataset 10 days ago

albertvillanova/tmp-state-on-load-ds

Preview • Updated 10 days ago • 5

updated a Space 10 days ago

Sleeping

🏆

Tmp State On Load

upvoted a collection 19 days ago

🤖 Agents

Collection

17 items • Updated 6 days ago • 35

upvoted a collection 22 days ago

SmolLM2

Collection

State-of-the-art compact LLMs for on-device applications: 1.7B, 360M, 135M • 15 items • Updated about 10 hours ago • 181

updated a Space 22 days ago

Running

🏆

Open LLM Leaderboard Model Comparator

Compare Open LLM Leaderboard results

posted an update 23 days ago

Post

1262

🚀 New feature of the Comparator of the 🤗 Open LLM Leaderboard: now compare models with their base versions & derivatives (finetunes, adapters, etc.). Perfect for tracking how adjustments affect performance & seeing innovations in action. Dive deeper into the leaderboard!

🛠️ Here's how to use it:
1. Select your model from the leaderboard.
2. Load its model tree.
3. Choose any base & derived models (adapters, finetunes, merges, quantizations) for comparison.
4. Press Load.
See side-by-side performance metrics instantly!

Ready to dive in? 🏆 Try the 🤗 Open LLM Leaderboard Comparator now! See how models stack up against their base versions and derivatives to understand fine-tuning and other adjustments. Easier model analysis for better insights! Check it out here: open-llm-leaderboard/comparator 🌐

posted an update 30 days ago

Post

3100

🚀 Exciting update! You can now compare multiple models side-by-side with the Hugging Face Open LLM Comparator! 📊

open-llm-leaderboard/comparator

Dive into multi-model evaluations, pinpoint the best model for your needs, and explore insights across top open LLMs all in one place. Ready to level up your model comparison game?

upvoted an article about 1 month ago

Article

Let's talk about LLM evaluation

•

May 23

• 134

upvoted 2 collections about 1 month ago

The Big Benchmarks Collection

Collection

Gathering benchmark spaces on the hub (beyond the Open LLM Leaderboard) • 13 items • Updated 10 days ago • 162

Open LLM Leaderboard best models ❤️‍🔥

Collection

A daily uploaded list of models with best evaluations on the LLM leaderboard: • 60 items • Updated about 1 hour ago • 446

posted an update about 1 month ago

Post

1217

🚨 Instruct-tuning impacts models differently across families! Qwen2.5-72B-Instruct excels on IFEval but struggles with MATH-Hard, while Llama-3.1-70B-Instruct avoids MATH performance loss! Why? Can they follow the format in examples? 📊 Compare models: open-llm-leaderboard/comparator

updated a Space about 1 month ago

Runtime error

🐨

Tmp Download

New activity in wikimedia/structured-wikipedia about 1 month ago

Dataset Viewer issue: TypeError: Couldn't cast array

#5 opened 2 months ago by

albertvillanova

upvoted an article about 1 month ago

Article

SmolLM - blazingly fast and remarkably powerful

Jul 16

• 271