LAiSER turns free text such as job postings and course syllabi into structured skills, knowledge and tasks, each matched to an entry in ESCO, O*NET, the UK Skills Classification or the Open Skills Network.
- About
- Architecture
- Requirements
- Setup and Installation
- Usage
- Try it in Google Colab
- Funding
- Authors
- Partners
LAiSER is a Python package for turning unstructured text about work and learning into structured, comparable skill data.
You give it a table of documents — job postings, course descriptions, syllabi. For each document, a language model extracts the skills it describes and, optionally, the knowledge and tasks behind them. LAiSER then matches every extracted phrase to its closest entries in established taxonomies using embedding similarity. The result is a table with one row per match: the phrase as written, the taxonomy entry it corresponds to, which taxonomy that entry comes from, and a similarity score.
Because results point at shared taxonomy entries rather than free-text phrases, they can be counted and compared across sources — the skills employers ask for can be set against the skills a program teaches, or tracked across thousands of postings.
The language model can be a hosted API (Gemini or OpenAI) or a model running on your own machine, including on CPU with no API key. Taxonomy alignment always runs locally against indexes bundled with the package.
LAiSER runs four stages for every document:
- Extraction — The text is placed into a prompt for its input type and sent to the language model, which returns candidate concepts.
- Parsing and deduplication — The response is parsed into phrases. Exact duplicates are dropped and near-duplicates are collapsed by embedding similarity.
- Taxonomy alignment — Each phrase is embedded and searched against bundled FAISS indexes of taxonomy entries. Matches above a per-type similarity threshold are kept.
- Output — Matches are returned as a single table, optionally with graph edges linking knowledge areas to the tasks they enable.
- Python
>=3.10; CI tests 3.10 through 3.13. - No GPU or API key is required. A GPU speeds up larger local models; hosted providers need their own API key.
- Supported model providers and their settings are listed in the documentation.
-
Install LAiSER from PyPI:
pip install laiser
-
Install with GPU extras:
pip install "laiser[gpu]" -
Install development dependencies from source:
pip install -e ".[dev]"
You can check if your machine has a GPU available with:
python -c "import torch; print(torch.cuda.is_available())"LAiSER is used as a Python package. The recommended API is SkillExtractorRefactored.
This runs a small open model locally on CPU. The model downloads on first use.
import pandas as pd
from laiser.skill_extractor_refactored import SkillExtractorRefactored
data = pd.DataFrame(
[
{
"Research ID": "job-001",
"description": "Build production machine learning systems in Python.",
}
]
)
extractor = SkillExtractorRefactored(model_id="Qwen/Qwen2.5-0.5B-Instruct", use_gpu=False)
results = extractor.extract_concepts(
data=data,
id_column="Research ID",
text_columns=["description"],
input_type="job_desc",
concepts=["skills"],
)
print(results[["Raw Concept", "Taxonomy Concept", "Taxonomy Source", "Correlation Coefficient"]])Small local models are convenient for trying LAiSER out; hosted or larger models extract more reliably.
import os
import pandas as pd
from laiser.skill_extractor_refactored import SkillExtractorRefactored
data = pd.DataFrame(
[
{
"Research ID": "job-001",
"description": "Build production machine learning systems in Python.",
}
]
)
extractor = SkillExtractorRefactored(
model_id="gemini",
api_key=os.getenv("GEMINI_API_KEY") or os.getenv("GOOGLE_API_KEY"),
use_gpu=False,
)
results = extractor.extract_concepts(
data=data,
id_column="Research ID",
text_columns=["description"],
input_type="job_desc",
concepts=["skills", "knowledge", "tasks"],
allowed_sources=["esco", "onet", "osn", "ukos"],
)
print(results.head())import os
import pandas as pd
from laiser.skill_extractor_refactored import SkillExtractorRefactored
data = pd.DataFrame(
[
{
"Research ID": "course-001",
"description": "Introduction to data visualization and exploratory analysis.",
"learning_outcomes": "Create dashboards, explain patterns in data, and evaluate charts.",
}
]
)
extractor = SkillExtractorRefactored(
model_id="gemini",
api_key=os.getenv("GEMINI_API_KEY") or os.getenv("GOOGLE_API_KEY"),
use_gpu=False,
)
results = extractor.extract_concepts(
data=data,
id_column="Research ID",
text_columns=["description", "learning_outcomes"],
input_type="course_syllabi",
concepts=["skills"],
allowed_sources=["esco", "onet", "osn", "ukos"],
)
print(results.head())| Option | Description |
|---|---|
model_id |
"gemini", "openai", or a Hugging Face model id to run locally |
api_key |
API key for hosted providers |
use_gpu |
run local models on GPU where available |
backend |
"llama_cpp" to run a local GGUF model |
temperature |
decoding temperature for every backend (default 0.0, greedy) |
seed |
seed for backends that accept one (default 42) |
concepts |
["skills"] (default), or add "knowledge" and "tasks" |
allowed_sources |
taxonomies to align against: "esco", "onet", "osn", "ukos" |
top_k |
maximum aligned matches per concept type, per document (default 25) |
return_edges |
return {nodes, edges} instead of only the results table |
output_csv_path |
also write the results to this CSV file |
Every option is described in the usage guide, and more snippets are in docs/examples.md.
Every backend decodes greedily by default (temperature 0.0) and uses a fixed seed where the backend accepts one, so repeated runs over the same input give the same result. To make results reproducible for others, pin every setting that affects them:
extractor = SkillExtractorRefactored(
model_id="Qwen/Qwen2.5-0.5B-Instruct", # pin the model
use_gpu=False,
temperature=0.0, # the default: greedy decoding
seed=42, # the default
)
results = extractor.extract_concepts(
data=data,
id_column="Research ID",
text_columns=["description"],
input_type="job_desc",
concepts=["skills"],
allowed_sources=["esco", "onet", "osn", "ukos"],
similarity_thresholds={"skill": 0.60, "knowledge": 0.50, "task": 0.50}, # the defaults
)When reporting results, also record the LAiSER version: it fixes the bundled taxonomy data and the embedding model used for alignment (sentence-transformers/all-MiniLM-L6-v2). Hosted providers do not guarantee identical output even at temperature 0.0. See Reproducibility for details.
The cookbook has complete analyses that open directly in Colab:
| Notebook | |
|---|---|
| Job skill analysis for job seekers | |
| University program skill analysis | |
| Pay equity analysis |






