Skip to content

language-learning-modelling/ll-datasets

View on GitHub ↗Related repositories →

active 2024-06-112024-11-25 (UTC)

Complete coverage26,622 / 26,622 hourly files (100%) · 2 absent upstream2023-08-152026-08-28 (UTC)
Events
51
Pushes
26
Pull requests
1
Issues
18
Stars
0
Forks
0

Activity over time

Daily event counts in the loaded window

Line chart, 168 days from 2024-06-11 to 2024-11-25. Pushes: 26 total, peak 5 in a day. Pull requests: 1 total, peak 1 in a day. Issues: 18 total, peak 5 in a day. Comments: 0 total, peak 0 in a day. Stars: 0 total, peak 0 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
berstearns452610

Recent activity

Latest issues, pull requests and releases

  • Issue#21berstearns2024-11-25 12:14
    create PODOL portuguse corpus downloader from gdrive
  • Issue#20berstearns2024-11-25 12:12
    create CEDEL2 downloader
  • Issue#19berstearns2024-11-25 12:12
    finish FCE downloader
  • Issue#18berstearns2024-11-21 09:40
    process celva-sp with teanga
  • Issue#17berstearns2024-11-11 12:00
    fix ll-datasets poetry install, all my clients are doing it to relative import not through poetry install
  • Issue#16berstearns2024-11-08 20:20
    look at the https://www.clarin.eu/resource-families/L2-corpora datasets and add to docs list
  • Issue#15berstearns2024-11-08 16:29
    process the COPOLE2 dataset
  • Issue#14berstearns2024-11-08 14:18
    add to docs the CEDEL2 and all the other related learner corpora to it. find the portuguese CEFR one, add ellipisis cefr dataset as well.
  • Issue#13berstearns2024-11-08 14:18
    process the CEDEL2 dataset
  • Issue#12berstearns2024-11-08 08:04
    create C4200m client
  • Issue#11berstearns2024-11-07 16:45
    create "run instance" abstraction
  • Issue#10berstearns2024-11-07 16:43
    add batch zlib option for the pipeline
  • Pull request#8berstearns2024-11-05 17:33
  • Issue#7berstearns2024-11-04 18:03
    refacotr project structure to clients have their own env and requirements, installing ll-datasets lib
  • Issue#6berstearns2024-11-04 17:51
    processing the pipeline with json.zlib output in each step
  • Issue#4berstearns2024-10-15 21:58
    make pipeline download, upload, resume work from remote data storage using rclone
  • Issue#3berstearns2024-10-15 21:57
    create option save as zlib batches
  • Issue#2berstearns2024-10-15 21:56
    implement initial normalization module
  • Issue#1berstearns2024-10-15 21:55
    abstract pandas_to_json + save_as_zlib code to be used by any Dataset class

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 0 stars here means stars gained during the window, not the repo's star count.