NVFP4 please!

#5
by Duonglv - opened

Hello,

Please quantize the NVFP4 version. It will help many people can try it. Because the FP16 is too big that can’t load in 32G Vram.

And does this model support multiplingual, such as Vietnamese?

Thank a lot.

Institute of Foundation Models org

We are trying it out. @aaryamonvikram

Our current model is mainly tested on English, so we don't know how well it performs on Vietnamese yet. We are having several programs to improve multilingual right now, especially our Jais series.

ModelOpt has Quantization Aware Distillation, should be this one, right?

Sign up or log in to comment