This website requires JavaScript.
0eb5c0c624
mac runners
Alex Cheema
2024-08-02 16:15:42 +01:00
e10aa581ac
cuda test
Alex Cheema
2024-08-02 16:06:54 +01:00
23d2432cb9
test cuda
Alex Cheema
2024-08-02 16:03:58 +01:00
fcdf57b873
t
Alex Cheema
2024-08-02 15:49:41 +01:00
f9c427467e
t
Alex Cheema
2024-08-02 15:45:54 +01:00
cf32ec9ab1
t
Alex Cheema
2024-08-02 15:44:07 +01:00
c5646ae00f
t
Alex Cheema
2024-08-02 15:42:59 +01:00
6b300112c8
t
Alex Cheema
2024-08-02 15:41:51 +01:00
4095dea805
t
Alex Cheema
2024-08-02 15:40:55 +01:00
9117125466
t
Alex Cheema
2024-08-02 15:40:08 +01:00
ec98b9cf65
t
Alex Cheema
2024-08-02 15:38:36 +01:00
fecd081047
t
Alex Cheema
2024-08-02 15:36:15 +01:00
49664c46c0
t
Alex Cheema
2024-08-02 15:35:25 +01:00
2564d7c2b4
docker runner
Alex Cheema
2024-08-02 15:34:11 +01:00
aa99c8ae50
trigger circleci
Alex Cheema
2024-08-02 15:09:27 +01:00
3313151010
test cuda
Alex Cheema
2024-08-02 14:58:19 +01:00
7599df7ffd
trigger circleci
Alex Cheema
2024-08-02 14:56:37 +01:00
0c3638f438
trigger circleci
Alex Cheema
2024-08-02 14:54:15 +01:00
a3b0650f30
use gpu.nvidia.medium
Alex Cheema
2024-08-02 14:49:48 +01:00
3a0c9c8582
linux-cuda image
Alex Cheema
2024-08-02 14:49:07 +01:00
af2f98bad4
run tinygrad tests on gpu.nvidia.small.gen2 (NVIDIA A10G 24GB)
Alex Cheema
2024-08-02 14:47:25 +01:00
0fd6bd91a0
run tinygrad with llama-3-8b
Alex Cheema
2024-08-02 14:20:24 +01:00
eafed8e16e
fix legacy model loading
Alex Cheema
2024-08-02 14:18:31 +01:00
08e8cacf01
run integration test for each inference engine
Alex Cheema
2024-08-02 13:47:06 +01:00
32bb44b3a2
request to both nodes in integration test, dont preload the model - exo should be robust against that
Alex Cheema
2024-08-02 11:51:53 +01:00
c06124c637
fix circleci badge
Alex Cheema
2024-08-02 11:48:52 +01:00
0b06516af8
Merge pull request #115 from exo-explore/circleci-project-setup
Alex Cheema and GitHub
2024-08-02 11:45:52 +01:00
f4b0f1ca02
replace tests badge with circleci
Alex Cheema
2024-08-02 11:45:09 +01:00
e3a524fd89
Merge pull request #114 from exo-explore/circleci-project-setup
Alex Cheema and GitHub
2024-08-02 11:41:53 +01:00
96969e3540
prefix matching logic
Alex Cheema
2024-08-02 11:28:48 +01:00
7c0923279c
trigger circleci
Alex Cheema
2024-08-02 11:25:34 +01:00
2a8f1ae4c5
get rid of caches. circeci downloads fast enough
Alex Cheema
2024-08-02 11:13:53 +01:00
54ea5dbb5c
only remove the matching prefix from the prompt if its length is less than the prompt
Alex Cheema
2024-08-02 11:11:16 +01:00
67ad3f57a1
use llama 3.1 in tests
Alex Cheema
2024-08-02 10:56:29 +01:00
92b66e275f
fix cache
Alex Cheema
2024-08-02 10:50:45 +01:00
cf3dae9570
circleci: separate hf, tinygrad caches
Alex Cheema
2024-08-02 10:45:53 +01:00
faadfa29dd
circleci chatgpt integration test
Alex Cheema
2024-08-02 10:44:35 +01:00
9a03991f1a
cache key
Alex Cheema
2024-08-02 10:39:46 +01:00
2065c7751a
orb
Alex Cheema
2024-08-02 10:27:12 +01:00
cf4cddccc1
port github workflow to circleci
Alex Cheema
2024-08-02 10:25:30 +01:00
21b0bf8777
Update config.yml
Alex Cheema and GitHub
2024-08-02 10:11:00 +01:00
997fcaff35
CircleCI Commit
Alex Cheema
2024-08-02 10:09:16 +01:00
f5755ea198
give explicit node ids when running on the same instance in tests, otherewise they use the same one because of sticky node ids
Alex Cheema
2024-08-01 01:24:50 +01:00
4faa6c06f4
add support for selective model downloading. related: #16
Alex Cheema
2024-08-01 01:22:41 +01:00
be3a09c28b
rm logs
Alex Cheema
2024-07-31 23:24:30 +01:00
0bfb8e3b6d
sticky node ids #16
Alex Cheema
2024-07-31 23:20:15 +01:00
980d5d2c5b
bring back TINYGRAD_DEBUG. not sure why it was removed
Alex Cheema
2024-07-31 22:56:53 +01:00
76766253cd
fix regression introduced by image_str for tinygrad
Alex Cheema
2024-07-31 22:53:46 +01:00
1d54f10514
pass on tinygrad set_on_download_progress
Alex Cheema
2024-07-31 22:50:04 +01:00
d6a7e46324
async model downloading with download progress. fixes #102 . related: #16 #104
Alex Cheema
2024-07-31 22:47:03 +01:00
5c67e24c35
smart prompt longest prefix matching to avoid sending the same text through the NN again. speeds up prefill significantly
Alex Cheema
2024-07-31 14:27:10 +01:00
94ac9463a7
fix model id for llama 3.1 405b now its finally on the hub
Alex Cheema
2024-07-31 10:16:30 +01:00
178fb75c84
fix image api prompt encoding
Alex Cheema
2024-07-30 23:56:30 +01:00
2d20000964
use AutoProcessor with use_fast=False since there's a bug with use_fast=True where whitespace is removed on single token decodes
Alex Cheema
2024-07-30 22:40:02 +01:00
0ec77e1a99
Merge pull request #88 from varshith15/main
Alex Cheema and GitHub
2024-07-30 20:02:47 +01:00
af1c7ce327
add support for image upload to tinychat for vision models
Alex Cheema
2024-07-30 20:01:35 +01:00
0d45a855fb
increase max request size to send raw images, make image download from url async, use chatgpt-compatible convention for images
Alex Cheema
2024-07-30 20:01:18 +01:00
e68d06f4ef
move model-selector styles to index.css
Alex Cheema
2024-07-30 14:51:22 +01:00
78db451d7e
add pillow to main dependencies
Alex Cheema
2024-07-30 14:27:45 +01:00
824f05263f
Merge branch 'main' into HEAD
Alex Cheema
2024-07-30 14:18:10 +01:00
142682645f
bump up tinygrad version
Alex Cheema
2024-07-30 10:02:20 +01:00
8d3d3df1dd
update readme
Varshith
2024-07-28 16:14:27 +05:30
acc94b50c7
chatgpt api integration
Varshith
2024-07-28 16:12:21 +05:30
33cbacf513
fix llava sanitize
Alex Cheema
2024-07-27 21:31:36 -07:00
2fb961fccd
stick to same convention as new llama
Alex Cheema
2024-07-27 21:04:13 -07:00
b44b917151
add pillow as testing dependency
Alex Cheema
2024-07-27 20:23:03 -07:00
2aa1e24ea9
remove unused torch import
Alex Cheema
2024-07-27 20:22:56 -07:00
833e7f3396
rename sharded_llava -> llava to match new convention
Alex Cheema
2024-07-27 20:19:55 -07:00
7d5eed1111
Merge branch 'main' into HEAD
Alex Cheema
2024-07-27 20:17:12 -07:00
044d189ccc
Merge pull request #94 from mzbac/mlx_refactor
Alex Cheema and GitHub
2024-07-27 20:02:51 -07:00
909d5ef8ba
Merge branch 'main' into mlx_refactor
Alex Cheema
2024-07-27 20:01:56 -07:00
63e51a8270
formatting
Alex Cheema
2024-07-27 20:00:39 -07:00
6695b019a2
format format.py
Alex Cheema
2024-07-27 19:57:18 -07:00
1dc08fecaa
increase max line length to 200
Alex Cheema
2024-07-27 19:56:50 -07:00
444137776a
formatting
Alex Cheema
2024-07-27 19:55:27 -07:00
a6bb8ddf41
update deepseek sanitize to shard layers first before handle switch
Anchen
2024-07-28 12:58:09 +10:00
cb217b7b77
format format.py
Alex Cheema
2024-07-27 19:57:18 -07:00
4cb36a7f55
increase max line length to 200
Alex Cheema
2024-07-27 19:56:50 -07:00
d94e3f9ce4
formatting
Alex Cheema
2024-07-27 19:55:27 -07:00
666b1c83ee
refactor(mlx): model sharding and add deepseek v2 support
Anchen
2024-07-28 12:46:38 +10:00
931ced7c01
fix a few more linter errors
Alex Cheema
2024-07-27 17:09:34 -07:00
57b2f2a4e2
fix ruff lint errors
Alex Cheema
2024-07-27 17:08:32 -07:00
ce761038ac
formatting / linting
Alex Cheema
2024-07-27 17:01:37 -07:00
f1bd5fe152
Merge pull request #90 from xeb/main
Alex Cheema and GitHub
2024-07-27 16:06:33 -07:00
f051ebe6e0
remove accidentally added files
Alex Cheema
2024-07-27 16:05:09 -07:00
5eafd5a305
try/except for decode, #75
Mark Kockerbeck
2024-07-27 13:00:14 -07:00
2849128d6a
processor load
Varshith
2024-07-28 01:03:34 +05:30
6ed76b3493
Merge branch 'main' into main
Varshith Bathini and GitHub
2024-07-28 01:00:08 +05:30
54993995dc
conflicts
Varshith
2024-07-28 00:44:41 +05:30
9d2616b9cf
shareded inference
Varshith
2024-07-28 00:30:34 +05:30
faa1319470
disable chatgpt api integration test, github changed something in their mac runners? perhaps time to switch over to circleci like mlx
Alex Cheema
2024-07-26 23:52:22 -07:00
67a1aaa823
check processes in github workflow
Alex Cheema
2024-07-26 23:37:11 -07:00
9a3ac273a9
Merge pull request #77 from Cloud1590/main
Alex Cheema and GitHub
2024-07-26 22:52:15 -07:00
628d8679b0
force mlx inference engine in github workflow, where it defaults to tinygrad because it's running on 'model': 'Apple Virtual Machine 1', 'chip': 'Apple M1 (Virtual)'
Alex Cheema
2024-07-26 21:46:41 -07:00
e856d7f7f9
log chatgpt integration test output from each process on github workflow failure
Alex Cheema
2024-07-26 21:37:42 -07:00
d2fa7b247e
Showing the message only if successfully decoded, #75
Mark Kockerbeck
2024-07-26 12:06:17 -07:00
f1cd5ae7a6
Merge branch 'main' of github.com:xeb/exo
Mark Kockerbeck
2024-07-26 12:04:18 -07:00
4f5ab78d9d
Addressing issue #75 to avoid decoding binary packets
Mark Kockerbeck
2024-07-26 12:03:49 -07:00
7cbf6a35bd
working test
Varshith
2024-07-26 19:12:42 +05:30
5a23376059
add log_request middleware if DEBUG>=2 to chatgpt api to debug api issues, default always to llama-3.1-8b
Alex Cheema
2024-07-25 20:33:26 -07:00