Antonio V. Franco

Antonio V. Franco

My work focuses on LLMs, SLMs, and agentic systems, with an emphasis on fine-tuning, quantization, distillation, and RAG.

For partnerships and projects, contact me by email: contact@antoniovfranco.com

You can read more about me on

  • GitHub
  • X
  • Hugging Face

Articles

2026

August

  • Quantizing an LLM Isn’t Enough: How to Prove That a 4-bit Model Still Deserves to Go Into Production Aug 29

July

  • The "DeepSeek V4 Flash + GLM-5.2" Combination Is Currently Enough for Me (and Will Probably Be Enough for You Too) Jul 24

June

  • GLM-5.2: I thought Sonnet 4.5 and similar open models were enough, but… Jun 25
  • Agentic AI, SLMs, and Why Models Above US$0.50 Output per 1M Tokens Are Equivalent to Burning Money Jun 13

2025

November

  • QDoRA Explained: Why It Became the New PEFT Standard in 2025 Nov 11

Antonio V. Franco

  • Antonio V. Franco
  • contact@antoniovfranco.com

Antonio V. Franco's personal blog featuring technical studies and work on LLMs, SLMs, fine-tuning, RAG, and agents.