diff --git a/README.md b/README.md index e7e647b..36e8f12 100644 --- a/README.md +++ b/README.md @@ -1,3 +1,192 @@ # fundi-scraper-api -Simple scrapter to turn all upcomming und past events at the Fundbureau into a REST API just because I am bored as fuck and want a small project while im in my hometown. Also this will be useful for me so I can ask my voice assistant for upcomming stuff. \ No newline at end of file +Simple scraper that turns all upcoming and past events at the +[Fundbureau](https://fundbureau.de/) (Hamburg) into a REST API — a small project +for a bored afternoon in my hometown, and handy for asking a voice assistant +what's coming up. + +Two parts: + +1. **`scraper/`** — a Playwright + BeautifulSoup script that scrapes the + Fundbureau site and writes the events to a JSON file. +2. **`api/`** — a FastAPI app that serves that JSON file as a REST API. + +``` +fundbureau.de ──scraper──▶ fundi-scraped-output.json ──api──▶ REST endpoints +``` + +## Requirements + +- Python 3.11+ +- For the scraper: `playwright`, `beautifulsoup4` (plus `playwright install chromium`) +- For the API: `fastapi`, `uvicorn` (see [`api/requirements.txt`](api/requirements.txt)) + +```bash +pip install beautifulsoup4 playwright +playwright install chromium +pip install -r api/requirements.txt +``` + +## 1. Scraper + +Located in [`scraper/scraper.py`](scraper/scraper.py). Run it from the repo root: + +```bash +python scraper/scraper.py +``` + +This writes `fundi-scraped-output.json` to the current directory. + +### Options + +| Flag | Description | Default | +|------|-------------|---------| +| `--url URL` | URL to scrape | `https://fundbureau.de/` | +| `--output-file NAME` | Output JSON file name | `fundi-scraped-output.json` | +| `--output-path PATH` | Directory for the output file | `./` | +| `--load-local-file` | Parse a local HTML file instead of scraping the live site (pass the file path as `--url`) | off | +| `--ignore-ticket-link` | Don't extract ticket links; leaves `event_ticket_link` empty and `event_free` null | off | + +### Output format + +The output is a JSON array of event objects: + +```json +[ + { + "event_date": "04.09.26", + "event_name": "Trance", + "event_artists": [ + [ + { "artist_name": "Skkin Velvet", "artist_play_time_start": "0", "artist_play_time_end": "Open end" }, + { "artist_name": "Bizarre", "artist_play_time_start": "23", "artist_play_time_end": "3" } + ] + ], + "event_starttime": "EINLASS 23:00", + "event_ticket_link": "https://www.ticketmaster.de/venue/fundbureau-hamburg-tickets/hamfundb/701", + "event_free": false + } +] +``` + +| Field | Type | Notes | +|-------|------|-------| +| `event_date` | string | `dd.mm.yy` | +| `event_name` | string | | +| `event_artists` | array of arrays of objects | outer array is the line-up grouping; each artist has `artist_name`, `artist_play_time_start`, `artist_play_time_end` (times are hours as strings, `"n/a"`, or `"Open end"`) | +| `event_starttime` | string | raw door-time text, e.g. `"EINLASS 23:00"` | +| `event_ticket_link` | string | URL, or `"n/a"` for free events, or `""` if unknown | +| `event_free` | bool \| null | `true` if admission is free, `null` when `--ignore-ticket-link` is used | + +## 2. REST API + +Located in [`api/`](api/). It loads a scraped JSON file into memory and exposes it +over HTTP. + +### Run + +```bash +cd api +uvicorn main:app --reload +``` + +- API root: http://127.0.0.1:8000 +- Interactive docs (Swagger UI): http://127.0.0.1:8000/docs +- OpenAPI schema: http://127.0.0.1:8000/openapi.json + +By default the API reads `fundi-scraped-output.json` from the repo root. Override +with the `DATA_FILE` environment variable: + +```bash +DATA_FILE=/path/to/events.json uvicorn main:app +``` + +### Endpoints + +| Method | Path | Description | +|--------|------|-------------| +| `GET` | `/health` | Service status, number of events loaded, resolved data file path | +| `GET` | `/events` | List events, with optional filters (below) | +| `GET` | `/events/{index}` | A single event by its 0-based position in the file; `404` if out of range | +| `GET` | `/artists` | Deduplicated, case-insensitively sorted list of all artist names | +| `POST` | `/reload` | Re-read the JSON file from disk (call after re-running the scraper); `500` if the file is missing | + +### `/events` query parameters + +All filters are optional and combine with AND. + +| Param | Type | Meaning | +|-------|------|---------| +| `free` | bool | `true` = only free events, `false` = only paid events | +| `name` | string | Case-insensitive substring match on `event_name` | +| `artist` | string | Case-insensitive substring match on any artist name in the line-up | +| `date` | string | Exact match on the raw `event_date` string (`dd.mm.yy`) | +| `upcoming` | bool | `true` = event date is today or later, `false` = past events (events with an unparseable date are excluded) | + +### Response shapes + +`GET /events` returns: + +```json +{ + "count": 2, + "events": [ { "event_date": "…", "event_name": "…", "...": "…" } ] +} +``` + +`GET /events/{index}` returns a single event object (same schema as the scraper +output). `GET /artists` returns a plain JSON array of strings. + +### Examples + +```bash +# All events +curl http://127.0.0.1:8000/events + +# Only free events +curl 'http://127.0.0.1:8000/events?free=true' + +# Upcoming events featuring an artist whose name contains "randali" +curl 'http://127.0.0.1:8000/events?artist=randali&upcoming=true' + +# Everything on a given night +curl 'http://127.0.0.1:8000/events?date=05.09.26' + +# First event in the file +curl http://127.0.0.1:8000/events/0 + +# All known artists +curl http://127.0.0.1:8000/artists + +# Refresh after re-scraping +python scraper/scraper.py && curl -X POST http://127.0.0.1:8000/reload +``` + +## Typical workflow + +```bash +# 1. Scrape +python scraper/scraper.py + +# 2. Serve +cd api && uvicorn main:app --reload + +# 3. (later) re-scrape and hot-reload the API without restarting it +python scraper/scraper.py +curl -X POST http://127.0.0.1:8000/reload +``` + +## Project layout + +``` +. +├── scraper/ +│ └── scraper.py # scrapes fundbureau.de -> JSON +├── api/ +│ ├── main.py # FastAPI app + routes +│ ├── datasource.py # loads/reloads the JSON, date & artist helpers +│ ├── models.py # Pydantic models for the event schema +│ ├── requirements.txt +│ └── README.md +└── fundi-scraped-output.json # example scraper output / default API datasource +```