When you’re building automation that needs to think, n8n’s visual workflow builder gets you 80% of the way there. The remaining 20% usually requires custom logic, LLM integration, or stateful processing that the standard nodes don’t handle. That’s where Python script nodes become essential. You get the speed and clarity of low-code workflows combined with the flexibility of real code.
I’ve built several production systems this way, and the pattern is consistent: start with what n8n gives you out of the box, then layer in Python for the parts that matter. Document processing pipelines, multi-step approval workflows, and intelligent data enrichment all follow the same principles.
Why Python in n8n?
n8n ships with JavaScript execution, which is fine for simple transformations. But Python handles several things better. Libraries like LangChain, Pydantic, and pandas are mature and widely used. If you’re already working in Python across your infrastructure, adding Python nodes to n8n means consistent logic, shared utilities, and easier knowledge transfer within your team.
More importantly, Python’s ecosystem for AI is simply ahead. If you need to call an LLM, process documents, or work with structured data validation, Python has the tools ready. You can write and test locally, then deploy the same code into a script node.
Setting Up Python Script Nodes
n8n executes Python through a dedicated runtime. You can either use n8n’s built-in Python support (if your instance has it enabled) or run a separate Python worker. For production, a dedicated worker is cleaner: it isolates dependencies, scales independently, and doesn’t compete for resources with the main n8n process.
Basic setup:
pip install n8n-python-worker
n8n-python-worker --port 5678
Then configure n8n to connect to this worker in your environment or settings. From there, any Python script node will execute against this runtime.
Inside a script node, you receive input data and return output:
import json
# Input comes as a list of items
# Each item has the data from the previous node
data = items[0]['json']
# Process it
result = {
'processed': True,
'input_count': len(data.get('records', [])),
'timestamp': str(datetime.now())
}
# Return as a list of items
return [{'json': result}]
The pattern is straightforward: receive items, transform them, return items. n8n handles the data flow between nodes.
Integrating LLMs
This is where Python nodes shine. You can call OpenAI, local models, or any LLM API from within a workflow step.
For OpenAI:
from openai import OpenAI
import json
client = OpenAI(api_key=os.getenv('OPENAI_API_KEY'))
text = items[0]['json']['text']
response = client.chat.completions.create(
model='gpt-4',
messages=[
{'role': 'system', 'content': 'You are a data classification assistant.'},
{'role': 'user', 'content': f'Classify this text: {text}'}
],
temperature=0.3
)
classification = response.choices[0].message.content
return [{
'json': {
'original_text': text,
'classification': classification
}
}]
For local models via Ollama or similar:
import requests
local_model_url = os.getenv('LOCAL_LLM_URL', 'http://localhost:11434')
text = items[0]['json']['text']
response = requests.post(
f'{local_model_url}/api/generate',
json={
'model': 'mistral',
'prompt': f'Extract entities from: {text}',
'stream': False
}
)
result = response.json()['response']
return [{'json': {'extracted': result}}]
The key difference: OpenAI costs money per token, but runs anywhere. Local models run free but require infrastructure. Most teams use both depending on the task complexity and volume.
The Hybrid Principle: Deterministic and AI Steps
Production workflows aren’t purely AI-driven. The most reliable patterns combine deterministic logic with AI steps, each handling what it does best. Deterministic steps handle explicit rules: formatting dates, validating email addresses, routing based on status fields. AI steps handle ambiguity: summarizing documents, classifying intent, extracting structured data from unstructured text.
This principle matters because it makes debugging easier. When something fails, you immediately know whether the issue is in your logic (predictable, easy to fix) or in the AI step (where you focus testing and iteration). As noted in production guidance from the n8n team, wrapping AI steps in deterministic logic or validation steps that handle predictability is the fix for reliability problems.
In practice, this means:
- Clean and validate data before it reaches your LLM
- Check that required fields exist and contain valid data
- Route decisions through explicit rules when the criteria are known
- Validate the AI output before it moves downstream
Handling State and Error Recovery
Workflows fail. Network timeouts, API rate limits, malformed data. A production pipeline needs to recover gracefully.
n8n provides error handling nodes, but Python script nodes can implement sophisticated retry logic:
import time
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
def requests_with_retries():
session = requests.Session()
retry = Retry(total=3, backoff_factor=0.5, status_forcelist=[429, 500, 502, 503, 504])
adapter = HTTPAdapter(max_retries=retry)
session.mount('http://', adapter)
session.mount('https://', adapter)
return session
data = items[0]['json']
session = requests_with_retries()
try:
response = session.post('https://api.example.com/process', json=data, timeout=10)
response.raise_for_status()
return [{'json': response.json()}]
except requests.exceptions.RequestException as e:
return [{'json': {'error': str(e), 'status': 'failed'}}]
For workflows that need to remember state across executions (like tracking which documents have been processed), store it externally: a database, Redis, or n8n’s built-in database if you’re using a managed instance.
import sqlite3
from datetime import datetime
db_path = os.getenv('STATE_DB_PATH', '/tmp/workflow_state.db')
conn = sqlite3.connect(db_path)
cursor = conn.cursor()
# Check if already processed
doc_id = items[0]['json']['document_id']
cursor.execute('SELECT processed_at FROM documents WHERE id = ?', (doc_id,))
row = cursor.fetchone()
if row:
return [{'json': {'status': 'already_processed', 'processed_at': row[0]}}]
# Process and mark as done
process_result = do_processing(items[0]['json'])
cursor.execute('INSERT INTO documents (id, processed_at) VALUES (?, ?)', (doc_id, datetime.now()))
conn.commit()
conn.close()
return [{'json': process_result}]
Real-World Example: Document Processing Pipeline
Let’s walk through a practical workflow: take uploaded documents, extract text, classify them, and route them to different teams.
The workflow has these steps:
- Webhook trigger receives a document upload
- Python node extracts text (using pdf2image, pytesseract, or similar)
- Python node calls an LLM to classify the document type
- Router node sends to different approval queues based on classification
- Error handler node logs failures to a database
The extraction node:
import PyPDF2
from io import BytesIO
import base64
pdf_base64 = items[0]['json']['file_content']
pdf_bytes = base64.b64decode(pdf_base64)
pdf_file = BytesIO(pdf_bytes)
reader = PyPDF2.PdfReader(pdf_file)
text = ''
for page in reader.pages:
text += page.extract_text()
return [{
'json': {
'document_name': items[0]['json']['file_name'],
'extracted_text': text,
'page_count': len(reader.pages)
}
}]
The classification node uses an LLM:
from openai import OpenAI
import json
client = OpenAI(api_key=os.getenv('OPENAI_API_KEY'))
text = items[0]['json']['extracted_text'][:2000] # Limit to first 2000 chars
response = client.chat.completions.create(
model='gpt-4',
messages=[
{
'role': 'system',
'content': 'Classify the document type. Return JSON with keys: type, confidence, summary.'
},
{
'role': 'user',
'content': f'Classify this document:
{text}'
}
],
temperature=0.2
)
classification = json.loads(response.choices[0].message.content)
return [{
'json': {
'document_name': items[0]['json']['document_name'],
'classification': classification,
'extracted_text': items[0]['json']['extracted_text']
}
}]
Then the router uses the classification to determine the next step. If it’s an invoice, route to finance. If it’s a contract, route to legal. And so on.
Scaling in Production
A single n8n instance with a single Python worker can handle thousands of workflows per day, depending on complexity. But when you need to scale, several patterns work well.
Run multiple Python workers and configure n8n to load-balance across them:
n8n-python-worker --port 5678
n8n-python-worker --port 5679
n8n-python-worker --port 5680
Then in n8n’s environment, list all workers:
PYTHON_FUNCTION_ALLOW_BUILTIN=true
PYTHON_FUNCTION_WORKER_URLS=http://localhost:5678,http://localhost:5679,http://localhost:5680
For CPU-heavy work like LLM calls or document processing, use n8n’s queue mode. Separate the execution from the webhook handler, so incoming requests don’t block waiting for slow operations.
Monitor execution times. If a Python script regularly takes more than 30 seconds, consider breaking it into smaller steps or running it asynchronously. n8n’s execution logs show you exactly where time is spent.
Keep Python dependencies minimal. Each script node needs to load its imports, so larger dependency trees slow execution. Use virtual environments to ensure only what you need is available.
Dependencies and Testing
In a Python worker, install dependencies once and they’re available to all scripts:
pip install openai pydantic requests PyPDF2 python-dotenv
For local testing, write your script as a normal Python function and test it before pasting into n8n:
def process_document(text):
from openai import OpenAI
client = OpenAI()
# ... LLM call ...
return result
# Test locally
test_result = process_document('Sample invoice text')
assert test_result['type'] == 'invoice'
Then adapt it to the n8n format (wrapping in the items loop) only when you’re confident it works.
Common Pitfalls in Python n8n Workflows
Long-running scripts without timeouts can hang n8n. Always set timeout values on external API calls. A 10 to 30 second timeout is typical for most integrations.
Logging to stdout works, but for production, send logs to a proper logger. n8n captures stdout but querying logs later is tedious. Use a structured logging library and send logs to a centralized service.
Hardcoding API keys is a security risk. Use environment variables and rotate them regularly. Store sensitive values in n8n’s credentials system or external secret management.
If a Python script works locally but fails in n8n, the issue is usually missing dependencies or environment variables. Check the execution logs carefully. The Python worker may have a different Python version or missing packages. Test the exact environment you’ll run in production.
Assuming all LLM outputs are valid JSON is a common mistake. Always parse and validate the response. Use structured output parsing if available, then add a validation step to check semantic correctness, not just structural correctness.
Wrapping Up
Python script nodes transform n8n from a workflow builder into a platform for intelligent automation. You get the visual clarity of low-code workflows, the flexibility of real code, and the power of Python’s AI ecosystem. For document processing, data enrichment, approval workflows, and any task that needs logic beyond standard integrations, this approach scales well.
Start small: add one Python node to an existing workflow. Test it, monitor it, then expand. The pattern is proven and the tooling is mature. You’re building enterprise automation without the overhead of traditional backend services.
Can I use Python in n8n without a separate worker?
Yes, n8n supports inline Python execution if the instance has Python installed and enabled. However, for production, a separate worker is recommended because it isolates dependencies, scales independently, and doesn’t compete for resources with the main n8n process.
What’s the difference between calling OpenAI and running a local LLM from n8n?
OpenAI via API is faster to set up and handles complex reasoning well, but incurs per-token costs and requires internet connectivity. Local models (like Mistral via Ollama) run free and offline, but require server infrastructure and typically handle simpler tasks. Most teams use both depending on task complexity and volume.
How do I handle API rate limits in Python script nodes?
Use the requests library with exponential backoff and retry logic. The HTTPAdapter with Retry strategy automatically retries failed requests with increasing delays. For external APIs with strict limits, implement a local cache or queue to batch requests and spread them over time.
Can I store workflow state between executions?
Yes. Store state in an external database (PostgreSQL, MongoDB, SQLite), a cache layer (Redis), or n8n’s built-in database if using a managed instance. Query the state at the start of a workflow, update it at the end. This works well for tracking processed documents, approval statuses, or any data that needs to persist across runs.
What happens if a Python script times out?
n8n will mark the execution as failed and trigger any error handlers you’ve configured. Always set explicit timeout values on external API calls (typically 10 to 30 seconds). For long-running processes, break them into smaller steps or run them asynchronously and poll for results.
How do I ensure my AI outputs are reliable?
Wrap AI steps in deterministic validation logic. Clean and validate data before it reaches your LLM. Use structured output parsing to enforce schema compliance. Add a validation step after parsing to check semantic correctness. Route low-confidence outputs to human review. This hybrid approach, combining deterministic and AI steps, is the foundation of reliable production workflows.