A web demo based on Gradio is provided in repo.
Support models list:
- ChatGLM
- ChatGLM2
- ChatGLM3
- ChatGLM4
- Llama2
- Llama3
- Gemma
- Yi
- Baichuan2
- Qwen
- Qwen2
Please refer to Installation. This example supports use source code which means you don't need install xFasterTransformer into pip and just build xFasterTransformer library, and it will search library in src directory.
Please refer to Prepare model
- Please refer to Prepare Environment to install oneCCL.
- Python dependencies.
PS: Due to the potential compatibility issues between the model file and the
pip install gradio transformers accelerate tiktoken transformers_stream_generator
transformersversion, please select the appropriatetransformersversion.
After the web server started, open the output URL in the browser to use the demo. Please specify the paths of model and tokenizer directory, and data type. transformer's tokenizer is used to encode and decode text so ${TOKEN_PATH} means the huggingface model directory.
# Recommend preloading `libiomp5.so` to get a better performance.
# or LD_PRELOAD=libiomp5.so manually, `libiomp5.so` file will be in `3rdparty/mkl/lib` directory after build xFasterTransformer.
export $(python -c 'import xfastertransformer as xft; print(xft.get_env())')
# run single instance like
python examples/web_demo/ChatGLM.py \
--dtype=bf16 \
--token_path=${TOKEN_PATH} \
--model_path=${MODEL_PATH}
# run multi-rank like
OMP_NUM_THREADS=48 mpirun \
-n 1 numactl -N 0 -m 0 python examples/web_demo/ChatGLM.py --dtype=bf16 --token_path=${TOKEN_PATH} --model_path=${MODEL_PATH}: \
-n 1 numactl -N 1 -m 1 python examples/web_demo/ChatGLM.py --dtype=bf16 --token_path=${TOKEN_PATH} --model_path=${MODEL_PATH}: Parameter options settings:
-h,--helpshow help message and exit.-t,--token_pathPath to tokenizer directory.-m,--model_pathPath to model directory.-d,--dtypeData type, default usingfp16, supports{fp16, bf16, int8, w8a8, int4, nf4, bf16_fp16, bf16_int8, bf16_w8a8,bf16_int4, bf16_nf4, w8a8_int8, w8a8_int4, w8a8_nf4}.
shell
python web_demo_api.py --url http://local:8000/v1 -m xft
Parameter options settings:
-h,--helpshow this help message and exit-u,--urlbase url likehttp://local:8000/v1-m,--modelmodel name-t,--tokenAPI token key-i,--ipgradio server ip, default0.0.0.0-p,--portgradio server port, default7860-s,--sharetureorfalse, whether to create a share link