I would have to go back and reread the paper to be sure, but FF layers are applied position-wise, meaning independently and in parallel on all input tokens/positions. Because of that, I could imagine contexts where the sequence dimension isn't relevant, i.e., for computational complexity.