Luke Salamone's Blog

https://lukesalamone.github.io

A technical blog exploring machine learning, deep learning architectures, data structures, and systems programming with in-depth implementations.

Luke Salamone's Blog

Entries

  • Paper Summary: The Matthew Effect in RL

    Learning to Solve Hard Problems in RL for LLMs by Never Giving Up discusses an approach for countering what the autho...

  • Semantic Search in Under 3MB

    This project is a continuation of my previous autoresearch project, which optimized a reranking model to be under 10M...

  • Teaching a Neural Net to Fight

    MathJax.Hub.Config({ tex2jax: { inlineMath: [['$','$'], ['\\(','\\)']], displayMath: [['$$','$$'], ['\[','\]']], proc...

  • Definitions of Model

    There are a lot of meanings of the term “model” in machine learning and machine learning-adjacent fields. Depending o...

  • Autoresearch

    Leaving the autoresearch loop going, the LLM was able to make 7.8% progress on the distillation task. I saw Andrej Ka...

  • Autoresearch

    Leaving the autoresearch loop going, the LLM was able to make 7.8% progress on the distillation task. I saw Andrej Ka...

  • Distilling Stockfish with One Billion Positions

    TLDR: I extracted fens and stockfish evaluations for 3.9 billion chess positions. I then trained a neural network on ...

  • Distilling Stockfish with One Billion Positions

    TLDR: I extracted fens and stockfish evaluations for 3.9 billion chess positions. I then trained a neural network on ...

  • Aesthetic Graph Pruning

    #mapContainer1, #mapContainer2 { color: #000; } #mapContainer1 button, #mapContainer2 button { padding: 10px 14px; bo...

  • Graph Topology and Battle Royale Mechanics

    #mapContainer1, #mapContainer2 { color: #000; } #mapContainer1 button, #mapContainer2 button { padding: 10px 14px; bo...

  • How well can Stockfish Estimate Itself?

    MathJax.Hub.Config({ tex2jax: { inlineMath: [['$','$'], ['\\(','\\)']], displayMath: [['$$','$$'], ['\[','\]']], proc...

  • Multi-Agent RL

    In reinforcement learning, an “agent” is an entity which can observe the world and take actions. Multi-agent setups h...

  • Optimal Ask

    Let’s say that you are selling N widgets and you need to determine a price for your widgets. There are N customers, e...

  • Can a Neural Net Learn It?

    A neural network is a type of function originally inspired by how the brain works. However, what makes it particularl...

  • Keep Summer Safe

    I recently built a small multi-agent simulation inspired by Rick and Morty. The setup is simple: The car must neutral...

  • Keep Summer Safe

    I recently built a small multi-agent simulation inspired by Rick and Morty. The setup is simple: The car must neutral...

  • Notes on Deepseek R1

    DeepSeek R1 is a large language model which employs test-time compute to generate a response. Unlike many past decode...

  • Notes on Deepseek R1

    DeepSeek R1 is a large language model which employs test-time compute to generate a response. Unlike many decoder-bas...

  • Space Is Really Big

    More than 30 earths could fit between the earth and the moon. Our elementary school models of the solar system really...

  • Space Is Really Big

    More than 30 earths could fit between the earth and the moon. Our elementary school models of the solar system really...