amokrov dtrawins commited on
Commit
21140c6
·
1 Parent(s): 49864d8

Update README.md (#1)

Browse files

- Update README.md (a771c90c1ea0aa1a35756be2c68505b1d3bfcca2)
- Update README.md (25fb0dd11bd0ff882936ef702d0f567fdfa77272)


Co-authored-by: Dariusz Trawinski <dtrawins@users.noreply.huggingface.co>

Files changed (1) hide show
  1. README.md +37 -0
README.md CHANGED
@@ -96,6 +96,43 @@ You can find more detaild usage examples in OpenVINO Notebooks:
96
  - [LLM](https://openvinotoolkit.github.io/openvino_notebooks/?search=LLM)
97
  - [RAG text generation](https://openvinotoolkit.github.io/openvino_notebooks/?search=RAG+system&tasks=Text+Generation)
98
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
99
  ## Limitations
100
 
101
  Check the original [model card](https://huggingface.co/Qwen/Qwen3-Coder-30B-A3B-Instruct) for limitations.
 
96
  - [LLM](https://openvinotoolkit.github.io/openvino_notebooks/?search=LLM)
97
  - [RAG text generation](https://openvinotoolkit.github.io/openvino_notebooks/?search=RAG+system&tasks=Text+Generation)
98
 
99
+ ## Running Model with OpenAI client and [OpenVINO Model Server](https://github.com/openvinotoolkit/model_server)
100
+
101
+ 1a. Deploy model on Windows using [binary package](https://docs.openvino.ai/ovms_baremetal):
102
+ ```
103
+ ovms.exe --rest_port 8000 --source_model OpenVINO/Qwen3-Coder-30B-A3B-Instruct-int4-ov --model_repository_path models --tool_parser qwen3coder --target_device GPU --cache_size 2 --task text_generation
104
+ ```
105
+ 1b. Deploy model in a Docker container:
106
+ ```
107
+ docker run -d --user $(id -u):$(id -g) --rm -p 8000:8000 -v $(pwd)/models:/models --device /dev/dri --group-add=$(stat -c "%g" /dev/dri/render* | head -n 1) openvino/model_server:latest-gpu \
108
+ --rest_port 8000 --model_repository_path /models --source_model OpenVINO/Qwen3-Coder-30B-A3B-Instruct-int4-ov --tool_parser qwen3coder --target_device GPU --task text_generation
109
+ ```
110
+ 2. Install the client library:
111
+
112
+ ```
113
+ pip install openai
114
+ ```
115
+ 3. Run the client:
116
+ ```
117
+ from openai import OpenAI
118
+ client = OpenAI(
119
+ base_url="http://localhost:8000/v3",
120
+ api_key="unused"
121
+ )
122
+
123
+ stream = client.chat.completions.create(
124
+ model="OpenVINO/Qwen3-Coder-30B-A3B-Instruct-int4-ov",
125
+ messages=[{"role": "user", "content": "Hello."}],
126
+ stream=True,
127
+ tools=[],
128
+ )
129
+ for chunk in stream:
130
+ if chunk.choices[0].delta.content is not None:
131
+ print(chunk.choices[0].delta.content, end="")
132
+ ```
133
+
134
+ Also check how to use this model in an agentic flow with function calling, as shown in the [agentic demo](https://docs.openvino.ai/ovms_batching).
135
+
136
  ## Limitations
137
 
138
  Check the original [model card](https://huggingface.co/Qwen/Qwen3-Coder-30B-A3B-Instruct) for limitations.