Skip to content

AliHaider0343/Automated-Text-Selection-for-Raw-Data-Annotation-

View on GitHub ↗Related repositories →

The paper tackles redundancy in user-generated data with an innovative text selection method, removing duplicates and using vectorized pruning. It enhances dataset quality, reduces redundancy, and improves vocabulary diversity, as shown by metrics like Type-Token Ratio and Herfindahl-Hirschman Index.

active 2024-01-042024-01-04 (UTC)

Complete coverage27,182 / 27,184 hourly files (100%) · 2 absent upstream2023-08-152026-09-20 (UTC)
Events
2
Pushes
0
Pull requests
0
Issues
0
Stars
0
Forks
0

Activity over time

Daily event counts in the loaded window

Line chart, 1 days from 2024-01-04 to 2024-01-04. Pushes: 0 total, peak 0 in a day. Pull requests: 0 total, peak 0 in a day. Issues: 0 total, peak 0 in a day. Comments: 0 total, peak 0 in a day. Stars: 0 total, peak 0 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

Nobody pushed, opened or commented here in the loaded window — this repo's activity is stars and forks only.

Recent activity

Latest issues, pull requests and releases

No issue or PR events — this repo's activity is pushes only.

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 0 stars here means stars gained during the window, not the repo's star count.