Quickstart¶
This walks through one complete extraction. Pick a model below — the first needs nothing but the package.
1. Prepare your input¶
LAiSER takes a pandas DataFrame with an ID column and one or more text columns:
import pandas as pd
data = pd.DataFrame(
[
{
"Research ID": "job-001",
"description": "Build production machine learning systems in Python. "
"Deploy models with Docker and monitor them on AWS.",
}
]
)
2. Create an extractor¶
Runs on CPU. The model downloads on first use.
Other options, including GPU and llama.cpp, are on Model providers.
3. Extract and align¶
results = extractor.extract_concepts(
data=data,
id_column="Research ID",
text_columns=["description"],
input_type="job_desc",
concepts=["skills"],
)
print(results[["Raw Concept", "Taxonomy Concept", "Taxonomy Source", "Correlation Coefficient"]])
4. Read the results¶
results is a DataFrame with one row per extracted phrase that matched a taxonomy entry:
| Column | Meaning |
|---|---|
Research ID |
the document the match came from |
Type |
skill, knowledge or task |
Raw Concept |
the phrase the model extracted from the text |
Taxonomy Concept |
the taxonomy entry it matched |
Taxonomy Description |
that entry's description |
Taxonomy Source |
which taxonomy: esco, onet, osn or ukos |
Source Url |
link to the entry, where the taxonomy provides one |
Correlation Coefficient |
similarity between phrase and entry; higher is closer |
Each phrase is matched to its single closest taxonomy entry and kept only if that match clears the similarity threshold, so a document never produces more rows than the model extracted phrases.
For the input above, the local 0.5B model on a laptop CPU returned:
| Type | Raw Concept | Taxonomy Concept | Taxonomy Source | Correlation Coefficient |
|---|---|---|---|---|
| skill | Python programming | Program in Python | ukos |
0.70 |
| skill | Machine learning | Use machine learning to create or improve solutions | ukos |
0.66 |
| skill | Docker | Docker | onet |
0.75 |
Decoding is greedy by default, so running this again returns the same rows — see Reproducibility. Different models extract different concepts: a hosted or larger model typically finds more, and this small model missed "AWS".
Next steps¶
- Extract knowledge and tasks too, or restrict taxonomies — see Usage.
- Run a full analysis in the browser with the Cookbook notebooks.