Skip to content

API reference

SkillExtractorRefactored is the public entry point. Construct it once with a model choice, then call extract_concepts on as many tables as you like; the model and taxonomy indexes are loaded once and reused.

This page is generated from the docstrings in laiser/skill_extractor_refactored.py.

SkillExtractorRefactored

SkillExtractorRefactored(model_id: Optional[str] = None, hf_token: Optional[str] = None, api_key: Optional[str] = None, use_gpu: Optional[bool] = None, backend: Optional[str] = None, temperature: float = DEFAULT_TEMPERATURE, seed: Optional[int] = GENERATION_SEED)

Refactored skill extractor with improved separation of concerns.

This class provides a clean interface while delegating specific responsibilities to appropriate service classes.

Initialize the skill extractor.

Parameters:

Name Type Description Default
model_id str

Model ID for the LLM

None
hf_token str

HuggingFace token for accessing gated repositories

None
api_key str

API key for external services (e.g., Gemini)

None
use_gpu bool

Whether to use GPU for model inference

None
backend str

Backend to use for LLM inference (e.g., "llama_cpp", "huggingface", "openai", "gemini")

None
temperature float

Decoding temperature for whichever backend is used. Defaults to 0.0 (greedy decoding), so repeated runs over the same input produce the same extraction. Raise it only when varied output is wanted.

DEFAULT_TEMPERATURE
seed int

Seed for backends that accept one: vLLM, llama.cpp, Gemini, and local Transformers models when temperature is above 0. Defaults to 42; pass None to leave it unset. The OpenAI API accepts no seed, so there reproducibility rests on greedy decoding alone.

GENERATION_SEED

extract_concepts

extract_concepts(data: DataFrame, id_column: str = 'Research ID', text_columns: List[str] = None, input_type: str = 'job_desc', top_k: Optional[int] = None, similarity_threshold: Optional[float] = None, levels: bool = False, batch_size: int = DEFAULT_BATCH_SIZE, warnings: bool = False, allowed_sources: Optional[List[str]] = None, concepts: List[str] = None, extract: List[str] = None, return_edges: bool = False, similarity_thresholds: Optional[Dict[str, float]] = None, timing: bool = False, output_csv_path: Optional[str] = None)

Extract concepts from text and align them to taxonomy data.

Parameters:

Name Type Description Default
data DataFrame

Input dataset

required
id_column str

Column name for document IDs

'Research ID'
text_columns List[str]

Column names containing text data

None
input_type str

Type of input data

'job_desc'
top_k int

Maximum number of aligned matches per concept type, per document (default: 25). With skills, knowledge and tasks all requested, one document can return up to three times this many rows.

None
similarity_threshold float

One minimum similarity score applied to every type. When omitted, the per-type defaults below apply.

None
similarity_thresholds dict

Per-type minimums, keys "skill", "knowledge", "task". Overrides similarity_threshold for the types it names. Defaults: {"skill": 0.60, "knowledge": 0.50, "task": 0.50}

None
levels bool

Accepted for compatibility; currently has no effect.

False
batch_size int

Accepted for compatibility; currently has no effect.

DEFAULT_BATCH_SIZE
warnings bool

Whether to show warnings

False
concepts list

Types to extract: "skills", "knowledge", "tasks", or ["all"]. Defaults to ["skills"].

None
extract list

Backward-compatible alias for concepts.

None
return_edges bool

If True, return {"nodes": pd.DataFrame, "edges": pd.DataFrame} where "edges" contains ENABLES edges (Knowledge → Task per skill). If False (default), return a plain pd.DataFrame.

False
timing bool

Accepted for compatibility with benchmark callers; currently has no effect.

False
output_csv_path str

If provided, write normalized results to this CSV path.

None

Returns:

Type Description
pd.DataFrame (when return_edges=False)

Normalized mixed-concept rows with: Research ID, Type, Raw Concept, Taxonomy Concept, Taxonomy Description, Taxonomy Source, Source Url, Correlation Coefficient

dict (when return_edges=True): {"nodes": pd.DataFrame, "edges": pd.DataFrame}

extract_and_align

extract_and_align(data: DataFrame, id_column: str = 'Research ID', text_columns: List[str] = None, input_type: str = 'job_desc', top_k: Optional[int] = None, similarity_threshold: Optional[float] = None, levels: bool = False, batch_size: int = DEFAULT_BATCH_SIZE, warnings: bool = False, allowed_sources: Optional[List[str]] = None, extract: List[str] = None, return_edges: bool = False, similarity_thresholds: Optional[Dict[str, float]] = None, timing: bool = False, output_csv_path: Optional[str] = None)

Backward-compatible skills-only extraction wrapper.

This method is retained to avoid breaking existing integrations that still call extract_and_align(...). It always runs the skills path. Use extract_concepts(...) for mixed concept extraction.