DandinPower
3a97a82f52
feat: add args to allow user can control the communication datatype
2025-08-23 13:43:14 +00:00
DandinPower
b780a35577
feat: refactor and fix bug for q8_0 and can easily support q4_0 and fp32 in future
2025-08-23 12:46:26 +00:00
DandinPower
2a10bca030
feat: add support for logging quantize and dequantize time
2025-08-23 19:17:22 +08:00
DandinPower
952a67f9a8
feat: refactor to make it can better support for future requirements (can switch using different communication type)
2025-08-23 17:54:04 +08:00
DandinPower
01f772b652
fix: fix the problem that zmq need pointer as input
2025-08-08 21:30:21 +08:00
DandinPower
ff310e7882
feat: turn the communication datatype from float32 to int8
2025-08-08 21:18:33 +08:00
DandinPower
7480ffbd30
Merge remote-tracking branch 'upstream/main' into merge/merge_86ca21e_from_upstream
2025-08-07 20:19:18 +08:00
DaninPower
96191d3b6f
feat: add CLI flag to control communication and computation logging
2025-08-02 12:55:12 +00:00
DandinPower
9bc57a10f6
feat: add tensor dumping functionality with shape information
...
- Add --dump-folder CLI argument to enable tensor dumping during network communication
- Implement binary dump format with tensor shape metadata (n_embed, n_tokens)
- Dump both send and receive tensors with unique filenames and counters
- Include proper parameter passing from CLI to llama_send_tensors/llama_recv_tensors functions
The dump format includes: element_type(1B) + n_embed(8B) + n_tokens(8B) + tensor_size(8B) + data
2025-07-28 09:57:07 +00:00
Li, Zonghang
bdf9d8e74b
llama-server: fix k-shift when output overlength
2025-07-17 21:03:41 +08:00
Li, Zonghang
b019a707b8
server: fix bugs
2025-07-13 13:42:24 +08:00
DandinPower
eb0cac1da5
fix: update the sbatch_tokens from %u to %lu to match unsigned long
2025-07-06 02:38:26 +08:00
DandinPower
adad23df61
feat: add more detail batch information
2025-07-05 21:40:58 +08:00
DandinPower
a399e49194
feat: add comm & compute log for further gantt chart analysis
2025-07-05 18:09:11 +08:00
DeEMO
2e8e42a5ad
Add speculative decoding support to the server and command-line interfaces
2025-06-30 09:35:35 +00:00
Zonghang Li
1ea2d61a97
speedup: add arg --keep-out-in-cuda to run the output layer on CUDA
2025-06-28 10:58:18 +04:00
Li, Zonghang
e8d3e5a631
update README
2025-06-27 20:16:30 +04:00
Zonghang Li
11ce0d58f7
fix compute buffer estimate: don't reverse CUDA VRAM for output layer
2025-06-27 12:42:16 +00:00
Li, Zonghang
aacfa8a231
fix compute buffer estimate: reserve 300 MiB VRAM to avoid potential OOM
2025-06-26 20:45:45 +04:00
Li, Zonghang
a05022c05a
communication: use barrier instead of manually adding delay
2025-06-26 17:30:47 +04:00
Li, Zonghang
3f27a25340
topo rebuild: add a delay to avoid packet interleaving
2025-06-26 14:50:58 +04:00
Li, Zonghang
729870fcd7
topo rebuild: add a delay to avoid packet interleaving
2025-06-26 14:47:34 +04:00
Li, Zonghang
72701ae872
fix compute buffer estimate: reserve 200 MiB VRAM to avoid potential OOM
2025-06-24 20:39:49 +04:00
Li, Zonghang
4dde8458cf
fix compute buffer estimate: reserve 100 MiB VRAM to avoid potential OOM
2025-06-24 19:29:10 +04:00
Li, Zonghang
90b1079d78
fix compute_buffer estimate: remove unused memory for CUDA device
2025-06-24 16:37:16 +04:00
Li, Zonghang
16ba3564ce
fix compute_buffer estimate: add context GPU usage
2025-06-24 16:09:59 +04:00
DandinPower
17edd6c96f
feat: remove the string_split function
2025-06-22 22:16:15 +08:00
Zonghang Li
45e8b0420c
fix compute buffer estimate: tested on cuda
2025-06-22 08:10:57 +00:00
Li, Zonghang
80e5b71b48
fix compute buffer estimate: tested on metal
2025-06-20 13:43:55 +04:00
Joseph Liaw
6a898ef9d8
docs: clarify split loading usage
2025-06-19 19:55:56 +08:00
Zonghang Li
dd589561b4
improve the computing buffer estimate
2025-06-19 08:02:43 +00:00
DeEMO
6ff38b2a0c
add args: data-port and signal-port
2025-06-17 12:00:04 +08:00
DeEMO
104e3b2356
fix: replace localhost to 127.0.0.1
2025-06-17 11:27:58 +08:00
Li, Zonghang
fbbc30c950
Merge branch 'speculative' into dev
2025-06-16 13:27:36 +04:00
Li, Zonghang
dc875bbef9
fix speculative decoding
2025-06-13 08:18:12 +04:00
DeEMO
d4618de991
fix: block when free socket
2025-06-12 12:26:10 +00:00
DeEMO
2039e3b0c1
fix: send and recv meta
2025-06-12 12:26:10 +00:00
DeEMO
d6c8d322cd
fix try_connect
2025-06-12 12:26:10 +00:00
DeEMO
d1b97f798e
support reconnection
2025-06-12 12:26:09 +00:00
Li, Zonghang
3e6d831930
fix seq_id mismatch between head and worker devices
2025-06-11 17:10:21 +04:00
DandinPower
1e7ae71ce5
feat: Add configurable network port options
2025-06-10 23:16:57 +08:00
Li, Zonghang
fb9b1f2b00
reformat llama.cpp
2025-06-09 13:04:22 +04:00
Li, Zonghang
22a6ddef13
fix batch decoding and dynamic batching
2025-06-07 00:53:56 +04:00
Lizonghang
e56be76bdf
assume only a single seq_id per token is needed
2025-06-07 00:42:44 +04:00
Lizonghang
d8aea899d1
fix n_seq_id and seq_id
2025-06-06 23:58:03 +04:00
Lizonghang
a1a2238831
add batch_all.n_seq_id and batch_all.seq_id to sync_meta
2025-06-06 23:36:53 +04:00
Lizonghang
68ecc8509d
add batch_all.logits to sync_meta
2025-06-06 22:58:48 +04:00
Lizonghang
500e066a2f
fix batch decoding and dynamic batching
2025-06-06 16:53:22 +04:00
Li, Zonghang
6439090920
reformat code
2025-06-03 23:53:24 +04:00
Li, Zonghang
a01fafd126
Merge branch 'main' into dev
2025-06-03 17:56:47 +04:00