Basal on public tenders: 2,452 notices for 72 cents
We took one week of the Polish public procurement bulletin (BZP), 21 to 27 September: 2,452 contract notices. We ran them through three basal‑1.5 models (mini, 4.5B and max), Bielik 11B as a plain language model, and a keyword filter. Basal picked the right industry more often and answered a single request several times faster. It missed more of the rare IT notices than Bielik, though, and the confidence thresholds from its calibration file need refitting.

We test Basal as a tender radar, a job usually done with a keyword filter: an IT or construction firm wants a daily feed of only the tenders it can win, and to miss none of the important ones.
The test: one week of notices, an answer key from CPV codes
The model gets the title and description of the contract, with the CPV codes removed from the text. CPV is the official procurement vocabulary, the buyer assigns the codes, and they make the answer key. The main code sets one of ten industries. A notice is for a construction firm if any of its codes starts with 45, and for an IT firm if one starts with 302, 48 or 72. We labelled nothing by hand, and we froze the questions and the test file before the first model run.
Each notice is one request with two questions about the same text. Here is one, in the layout of the examples on the Basal website.
- Example input: notice 2026/BZP 00453634/01 of 24 September, a delivery of disk arrays, Fibre Channel switches and a router.
- Questions to the model: “Do której branży należy to zamówienie?” (which industry, type
choice) and “Dla jakich firm jest to zamówienie?” (which firms, typemulti). - Defined answers: ten industries and two kinds of firm, each option with a description. The keys are hidden (
option_keys: hide), so the model reads only the Polish descriptions. - What the application does: a notice above the confidence threshold goes straight to the right firm's inbox, and a person reviews the rest.
JSON
{
"state": "Nazwa zamówienia: Jednorazowa dostawa infrastruktury sieciowej i pamięci masowej z podziałem na 2 części\nOpis: Część 1: Dostawa macierzy dyskowych i przełączników Fibre Channel wraz z licencjami i wdrożeniem\nCzęść 2: Dostawa przełącznika sieciowego / routera",
"questions": {
"industry": {
"type": "choice",
"instructions": "Do której branży należy to zamówienie?",
"criteria": {
"construction": "Roboty budowlane: budowa, remont, przebudowa obiektów, dróg, sieci i instalacji",
"engineering": "Projektowanie i usługi inżynieryjne: dokumentacja projektowa, nadzór inwestorski, geodezja",
"medical": "Medycyna: sprzęt i wyroby medyczne, leki, usługi zdrowotne",
"vehicles": "Pojazdy i transport: samochody, autobusy, części, usługi przewozowe",
"it": "IT: sprzęt komputerowy, oprogramowanie, usługi informatyczne",
"food": "Żywność i gastronomia: artykuły spożywcze, catering, usługi hotelowe i restauracyjne",
"energy": "Paliwa i energia: paliwa, energia elektryczna, gaz, opał",
"waste": "Odpady, sprzątanie i zieleń: odbiór odpadów, utrzymanie czystości, tereny zielone",
"furniture": "Meble i wyposażenie: meble, wyposażenie wnętrz, sprzęt gospodarstwa domowego",
"other": "Inne: zamówienie nie pasuje do żadnej z pozostałych grup"
},
"option_keys": "hide"
},
"relevance": {
"type": "multi",
"instructions": "Dla jakich firm jest to zamówienie?",
"criteria": {
"it": "Firma informatyczna: oprogramowanie, usługi IT, sprzęt komputerowy",
"construction": "Firma budowlana: roboty budowlane, remonty, instalacje budowlane"
},
"option_keys": "hide",
"threshold": 0.5,
"min": 0,
"max": 2
}
}
}POST /v1/systemone, exactly as we sent it to the basal v1.5.0 server. The model reads Polish, so the questions and options are in Polish.JSON
{
"answers": {
"industry": {
"type": "choice",
"choice": "it",
"confidence": 0.991
},
"relevance": {
"type": "multi",
"selected": ["it"],
"probabilities": {
"it": 0.823,
"construction": 0.0005
}
}
}
}model and usage fields, the full industry distribution, and threshold and set_confidence of the multi answer. We rebuilt selected from the probabilities by the engine's rule; the other values are from our run, rounded.Max chose IT with confidence 0.991, above the threshold in its calibration file (0.947), so the application would pass the notice on without a person. Mini, 4.5B and Bielik also chose IT. The keyword filter put it under “other”, because the title contains none of the words on its list.
Industry: every Basal ahead of Bielik
All three Basal models beat Bielik 11B, the model max was fine-tuned from, even mini, a 3.2 GB download. The 95% confidence intervals do not overlap: 77.8 to 80.9% for max, 71.5 to 75.0% for Bielik.
Some of the misses are a matter of interpretation: many notices carry codes from several industries, and the question allows one. If any industry among a notice's codes counts as a hit, max scores 84.9% and Bielik 78.8%.
Time and cost
Each model first got 200 notices one at a time, which measures response time, and then the whole week with 16 requests in flight, which measures throughput. Bielik has to generate a JSON answer, while Basal only reads the probabilities of the options, so it is several times faster on a single request.
With 16 requests in flight Bielik catches up: vLLM batches its requests and gets through 9.5 notices per second, almost twice as many as max (5.1) and almost as many as 4.5B (10.4). Mini: 28.5. At $0.583 an hour, a thousand notices cost from 0.6 cents on mini to 3.1 cents on max, so a week of BZP costs under eight cents on any model. The whole run, with installation, four model downloads and compilation, cost us $0.72.
Who the notice is for: construction is easy, IT is hard
The second question decides what lands in a firm's inbox. That week 932 notices had a construction code and only 128 an IT code.
- Construction. Max found 94% of these notices, and 96% of its flags were right. Bielik found 98.5%, but only 82% of its flags were right, which means 206 unneeded notices in the week. The keyword filter: 80% found and 77% of flags right.
- IT. Bielik found 88% of the notices, but only 54% of its flags were right. Mini found 59% with 82% of flags right, max 50% with 81%, the keyword filter 62% with 79%. Bielik was asked “Czy to zamówienie może zrealizować firma informatyczna…?” (can an IT firm deliver this contract), Basal the engine's template “Czy dotyczy: …?” (does this apply).
- The 4.5B model almost always answered “no” to the IT question and found 8 notices out of 128, although it picked IT as the industry for 80% of them.
For a rare label, a yes or no question with a 0.5 threshold is not enough on its own. When the IT signal was the answer to the industry question, each of the three Basal models found 78 to 80% of the IT notices with 66 to 73% of flags right. Another option is a separate threshold per label, fitted on your own data.
Fit the confidence threshold to your own data
Each Basal model ships a calibration file with a confidence threshold chosen on the author's data so that at most 1% of the decisions above it are wrong. On our industry question the same thresholds accepted 31% of notices at 3.3% error for 4.5B, 58% at 5.0% for max and 44% at 8.1% for mini. Even counting any industry among a notice's codes as a hit, the error is 2.1 to 5.7%. The engine's README says so plainly: with your own data, refit the threshold on a few hundred labelled requests.
A tender radar on Basal, step by step
- 01Take a few hundred notices from outside the week you will test on; the CPV codes give you the answer key. Fit the confidence threshold on them and send everything below it to a person.
- 02Start with basal‑1.5‑mini: 29 ms per notice and better than Bielik 11B on industry. Max adds 3.8 points but answers almost five times slower.
- 03Route on a
choicequestion whose options have descriptions. Check amultilabel on your own data before you trust it.
Sources
- 01Biuletyn Zamówień Publicznych, ogłoszenia o zamówieniunotices published 21 September 2026 to 27 September 2026
- 02Rozporządzenie Komisji (WE) nr 213/2008 w sprawie Wspólnego Słownika Zamówień (CPV)published 15 March 2008
- 03basal.si5.pl, Zastosowania: 18 scenariuszy z zapytaniem i odpowiedzią JSONpage as of 6 October 2026
- 04rkinas/basal, silnik v1.5.0 i READMEtag of 4 October 2026
- 05Hugging Face, Remek/basal‑1.5‑minirevision of 3 October 2026
- 06Hugging Face, Remek/basal‑1.5‑4.5Brevision of 3 October 2026
- 07Hugging Face, Remek/basal‑1.5‑maxrevision of 3 October 2026
- 08Hugging Face, speakleash/Bielik‑PL‑11B‑v3.0‑Instructrevision of 14 April 2026
- 09vLLM v0.30.0, informacja o wydaniupublished 22 September 2026
