This is called (multitrack) music transcription. There are some commercial solutions (AudioScore, AnthemScore, ...).
For OSS, look at Omnizart [1] and magenta/mt3 [2].
I suppose these models are trained on western / pop music, so they may not work nicely on ethnic music.
[1] https://github.com/Music-and-Culture-Technology-Lab/omnizart [2] https://github.com/magenta/mt3