I 100% agree the model is basically useless. that was never the deliverable. the artifact here is the inference engine, not the model living in it. 3.16M params at character level is just what fits in ~3MB of on-chip SRAM.
Plus I treated it as a good learning experience to get better with FPGA's but also system design.