Skip to content

burakaktas35/Accelerating-LLM-Inference-with-llama-cpp-python-and-.gguf-Files-for-Optimized-Performance

View on GitHub ↗Related repositories →

This project demonstrates the performance gains achieved by using llama-cpp-python with .gguf files compared to .safetensors with the Transformers library for LLM inference. The use of llama-cpp-python significantly decreases inference time, optimizing the speed of large language model completions in a local setting.

active 2024-11-182024-11-18 (UTC)

Complete coverage27,187 / 27,189 hourly files (100%) · 2 absent upstream2023-08-152026-09-20 (UTC)
Events
3
Pushes
1
Pull requests
0
Issues
0
Stars
0
Forks
0

Activity over time

Daily event counts in the loaded window

Line chart, 1 days from 2024-11-18 to 2024-11-18. Pushes: 1 total, peak 1 in a day. Pull requests: 0 total, peak 0 in a day. Issues: 0 total, peak 0 in a day. Comments: 0 total, peak 0 in a day. Stars: 0 total, peak 0 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
burakaktas351100

Recent activity

Latest issues, pull requests and releases

No issue or PR events — this repo's activity is pushes only.

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 0 stars here means stars gained during the window, not the repo's star count.