8 comments

  • dahnhiller 8 hours ago
    Hi everyone, I’m a friend of Mike’s; he’s having issues replying to the post at the moment, but hopes to post a thorough reply to the comments as soon as possible
  • stevefan1999 9 hours ago
    But the problem is not that your model is fast.

    Sure, you can go ASIC and go even faster, but the thing around GPU is that they scale well for both training and inference, and the technical floor is low.

    The level to get into FPGA design is insanely high, you've got to read timing diagrams, you need to know combinatorial and sequential logics and good sense of boolean algebra, you need to have an asynchronous signal based mindset which is vastly different from CPU/GPU, you need to know netlist and you need to endure the time it takes for the EDA to finish generating it. Yosys is still years behind Xilinx

    There is a reason GPUs are called accelerators; it sacrifices and does not try to really specialize on one particular thing, except high parallel dataflow and branch-free calculation. Otherwise we will all be using DSPs

    • IshKebab 44 minutes ago
      FPGAs aren't that difficult to program. Waves / timing diagrams are trivial to understand, they can just be very tedious to read. The different execution model to CPUs also isn't very hard to understand IMO. Hardware people love to say that it is, but it really isn't.

      The hardest bit is probably SystemVerilog - it's just such a terrible language for hardware design. Full of footguns and gotchas and weird limitations and undocumented or tool-dependent stuff that you have to just know (like what is synthesizable).

      Approximately nobody uses Yosys.

    • cgyvbunji 9 hours ago
      FPGAs are not power efficient at all vs GPUs and ASICs anyway, which is going to be especially true when they are fully saturated by LLM inference.
      • stevefan1999 8 hours ago
        That said, FPGA do provide a middle ground, but using it for speed and power efficient is not a forte, and the true value exactly comes from this focus alone: it allows you do emulate systhesis and verify that your logic is correct before you do full ASIC tapeout, e.g. building softcores for CPU validation

        Anything else is added and unintentional benefits.

  • haeseong 12 hours ago
    I didn't expect the 2,000 connection sweep to stay flat, since all of them are sharing one stream. What does per user latency look like at that end of the sweep?
    • all2 8 hours ago
      For those that don't have dead comments showing in their HN UI, Mike has a response adjacent to this one [0].

      [0] https://news.ycombinator.com/item?id=49244312

      @dang, the creator of this idea is having his comments killed off for some reason.

    • mikeayles 11 hours ago
      You're correct, the flat line is aggregate only. the fabric is saturated from a few dozen active clients onward, so extra connections can't buy throughput, they just queue. per-user p50/p95 across that same sweep: 17ms/30ms solo, 450ms/545ms at 100, ~2s/2.4s at 500, 3.8s/4.4s at 1000, 6.3s/9.4s at 2000. zero errors or drops at every stage. It degrades as a well-behaved queue, not a cliff, but nobody would call 6s at the top end snappy.

      The benchmark sweep is 2,000 concurrent active requesters hammering it constantly, real traffic is mostly lurkers, which cost a file descriptor and nothing else. the interactive feel actually gives out earlier than the queue math. The speculative-typing UI wants sub-second replies, and that budget blows around 100–150 simultaneous typists.

      I've been logging the stats since it went live, unfortunately it didn't hit FP. Peak was 10 concurrent connections (13 uniques in the busiest half hour), ~580 requests and ~37k tokens served, and at no point did two people actually have an inference in flight at the same moment which would have been the real test for the queue, every visitor got the fabric to themselves, p50 ~23ms. so the 2,000-conn drill was not stressed today. the one blemish: a single window with p95 ~57s, which lines up with the model-rotation FPGA reconfigure rather than load. A request that arrives mid-reflash waits out the ~25s swap. if this thread sends 50× more people, the queue math above says it holds.

      I need to discard the requests that overlap the model changeover for a truer result.

      • useiris 5 hours ago
        does the reflash actually stall every live connection, or just the ones whose request lands during that window? if the whole board goes dark for the full ~25s while any request is queued behind it, you could probably hide most of that behind partial reconfiguration, reflashing only the region holding the model weights while the sequencer and I/O logic on the rest of the fabric stay live and keep draining the queue. that's obviously a much bigger lift than what you've built here, but it would turn a hard stop into something closer to a brief latency bump for whoever's unlucky enough to hit it, rather than a shared 25s wall for everyone behind them in line.
  • serf 1 hour ago
    conceptually it's a cool idea.

    practically the results seem about as coherent as

      import random; print(random.choice(list(my_dict)))
    
    ..but way slower

    is there a practical use to a model this small?

  • peter_d_sherman 8 hours ago
    Ignore the naysayers!

    Any article, even the really good ones on HN, while they get positive comments, for whatever reason, always get a lot of negative ones, too...

    That is, the negative comments are absolutely unavoidable, even for people accomplishing great things!

    I personally think that what you've done is brilliant, absolutely brilliant!

    I can't wait to see more in this space...

    Brilliant, absolutely brilliant!

    • RetroTechie 6 hours ago
      Very nice indeed. Model weights in RAM blocks distributed all over a big FPGA: should be super helpful at minimizing RAM bandwidth bottlenecks. To say nothing of latency.

      But model(s) implemented are clearly too small to be useful as a 'chat partner'. Tried a couple of sentences - replies is just some gibberish coming out.

      This really needs a bigger FPGA, or some other application(s) where a tiny LLM does actually useful work. Barring that, generated tokens/sec is kind of a meaningless measure imho.

    • M4R5H4LL 8 hours ago
      I am also very cautious with people who tell me something impossible when I can trust my engineering skills and get a good sense that there is potentially a good outcome. In my experience, it simply means they don’t know how to do it, or are frustrated they couldn’t do it themselves and get into the spotlight.
  • mrheosuper 26 minutes ago
    [dead]
  • mikeayles 13 hours ago
    I started this about 10 weeks ago when the Taalas chatjimmy demo first did the rounds, Llama 8B baked into custom silicon, 17k tok/s for a single user. Their whole thesis is that inference is bound by reading the weights, so stop fetching them from far away. I wanted to see how far that idea stretches on a 'consumer hardware': every weight resident in SRAM, zero DRAM My chip only gives you ~3 MB to live in, so the model is 3.16M parameters and the tokens are characters, but the physics is the same bet. Only difference is their model is actually good!

    The annoying part: I wanted this live last week, which would have landed the same day as the AMD acquisition news, on AMD's own silicon no less. Life got in the way, so here it is a week late instead.

  • threadsnoop 12 hours ago
    [dead]