Tabby collected the source code from the related repository and built them into a dataset. I already have some proof-of-concept pipelines built [1] and [2], but I still need some time to polish the data pipeline..
[1]: https://github.com/TabbyML/tabby/blob/main/tabby/tools/repos...
[2]: https://github.com/TabbyML/tabby/blob/main/tabby/tools/train...