Ok so I’ve now tried three different MLX implementations for a local inference engine and all of them collapse when handed more than a few thousand tokens to <1 token/sec. All in the server wrapper though so I see some potential for making a native swift version to resolve this..
0
Replies
1
Boost
No replies yet
Be the first to share your thoughts.