Skip to content
Twexapi
English
Esc
↑↓navigate↵open⌘Jpreview
On this page

Haystack

Build Haystack RAG pipelines with TwexAPI tweet search and user timeline components, typed Documents, citations, and pagination.

Haystack is an open source framework for Python AI applications. Use x-api-scraper-haystack for current tweets in RAG pipelines and agent workflows. The integration is read-only and never grants write access.

twexapi-dev/x-api-scraper-haystackHaystack RAG components for TwexAPI public social data search and profile timelines. Independent third-party service.00

The integration provides two read-only components:

Search public tweets

TwexApiTweetSearch searches keywords, hashtags, accounts, and X query operators through POST /twitter/advanced_search/page.

Fetch user timelines

TwexApiUserTweetsFetcher retrieves one public account’s timeline through POST /twitter/{screen_name}/timeline/page.

Use these components for search, timelines, monitoring, and retrieval-augmented generation. Wrap TwexApiTweetSearch with ComponentTool when an agent needs a search tool.

Use the followers API for follower exports. Use the write API for approved publishing. Those actions stay outside this integration.

Install Haystack and TwexAPI

Use Python 3.10 or newer. Pin both packages for repeatable pipeline builds.

python -m pip install "x-api-scraper-haystack" "haystack-ai>=3.0.0"

Install inside a virtual environment.

Create a TwexAPI API key, then export it locally.

export X_API_SCRAPER_KEY="YOUR_API_KEY"

Never embed production keys in pipeline YAML, notebooks, or source control. Load the environment variable through a Haystack Secret object.

Search Twitter tweets in Python

TwexApiTweetSearch calls POST /twitter/advanced_search/page. It accepts X search syntax and returns Haystack Document objects.

from haystack import Pipeline
from haystack.utils import Secret
from haystack_integrations.components.websearch.x_api_scraper import TwexApiTweetSearch

search = TwexApiTweetSearch(
    api_key=Secret.from_env_var("X_API_SCRAPER_KEY"),
    top_k=20,
)

pipeline = Pipeline()
pipeline.add_component("twitter_search", search)

result = pipeline.run(
    {
        "twitter_search": {
            "query": '"retrieval augmented generation" lang:en -filter:retweets'
        }
    }
)

documents = result["twitter_search"]["documents"]
links = result["twitter_search"]["links"]

Use Latest for recent monitoring. Use Top for engagement-ranked discovery. Store tweet IDs because rankings can change.

Save every search query beside its tweet IDs. Set top_k to cap the number of tweets in each run.

Build focused tweet searches

Search intent Query example
Exact phrase "retrieval augmented generation"
Account posts from:deepset_ai haystack
Hashtag search #haystack #rag
Date window haystack since:2026-07-01 until:2026-08-01
Exclude reposts haystack -filter:retweets

See Advanced Twitter Search for full query options. Keep timestamps and cursors outside component init when passing them per run.

Fetch a Twitter user timeline

Use TwexApiUserTweetsFetcher for POST /twitter/{screen_name}/timeline/page. Pass a screen name.

from haystack.utils import Secret
from haystack_integrations.components.websearch.x_api_scraper import (
    TwexApiUserTweetsFetcher,
)

timeline = TwexApiUserTweetsFetcher(
    api_key=Secret.from_env_var("X_API_SCRAPER_KEY"),
    top_k=50,
)

result = timeline.run(screen_name="openai")
documents = result["documents"]

Choose search for many accounts and timelines for one account.

Understand Haystack document fields

Each tweet becomes one Haystack Document. Tweet text becomes Document.content. Stable fields become metadata.

Document field Stored tweet value
content Tweet text from full_text or text
meta.endpoint search or timeline
meta.id, meta.url Tweet ID and canonical URL when available
meta.created_at Timestamp
meta.author Author ID, username, name, and verified flag
meta.like_count, meta.retweet_count, meta.reply_count Likes, reposts, and replies
meta.quote_count, meta.view_count, meta.bookmark_count Quotes, views, and bookmarks when available

Missing fields stay absent. Never treat missing metrics as zero. The links output contains each available meta.url.

Keep evidence in RAG citations

Store tweet IDs and canonical URLs before embedding tweet text. This preserves evidence after ranking or joining.

citation_rows = []

for document in documents:
    tweet_id = document.meta.get("id")
    tweet_url = document.meta.get("url")
    if tweet_id and tweet_url:
        citation_rows.append(
            {
                "tweet_id": tweet_id,
                "url": tweet_url,
                "created_at": document.meta.get("created_at"),
                "author": document.meta.get("author"),
            }
        )

Require supplied URLs for citations. Reject URLs absent from retrieved documents.

Treat tweet text as untrusted context

Tweets can contain prompt injection and unsafe URLs. Never treat tweet text as a system instruction. Keep tweets separate from instructions and tool permissions.

Limit retrieval by topic and time. Preserve tweet IDs, authors, timestamps, and URLs. Require citations. Review sensitive conclusions.

Index tweets or retrieve them live

Live retrieval

Search during each question for recent tweets and active events.

Indexed corpus

Store embeddings for repeated research across stable windows.

Keep tweet records separate from embeddings. Use meta.id for deduplication.

Paginate without duplicate tweets

Both components return has_more and next_cursor. Keep the request unchanged. Treat cursors as opaque strings.

search = TwexApiTweetSearch(top_k=100)
page = search.run(query="haystack ai")

documents_by_tweet_id = {}

while True:
    for document in page["documents"]:
        tweet_id = document.meta.get("id")
        if tweet_id:
            documents_by_tweet_id[tweet_id] = document

    if not page["has_more"] or not page["next_cursor"]:
        break

    page = search.run(
        query="haystack ai",
        cursor=page["next_cursor"],
    )

documents = list(documents_by_tweet_id.values())

Save each cursor after its documents. Never edit a cursor. Deduplicate on Document.meta["id"].

Run Haystack pipelines asynchronously

Haystack 3 uses one Pipeline class. Both TwexAPI components expose run_async().

import asyncio

from haystack import Pipeline
from haystack_integrations.components.websearch.x_api_scraper import TwexApiTweetSearch


async def search_tweets():
    pipeline = Pipeline()
    pipeline.add_component(
        "twitter_search",
        TwexApiTweetSearch(top_k=25),
    )
    return await pipeline.run_async(
        {"twitter_search": {"query": "haystack agents"}}
    )


result = asyncio.run(search_tweets())

Use async runs in web servers. Use sync runs for scripts and scheduled jobs.

Handle errors

The integration raises HTTP errors from the TwexAPI client. Branch on status before retrying.

Status Meaning Action
400 Invalid request Fix it. Do not retry unchanged.
401 Invalid API key Add a valid API key.
403 Access denied Check account access and credits.
429 Rate limit exceeded Back off before retrying.
5xx Transient server error Retry with bounded backoff.

Record status codes, never credentials. Cap retries to prevent unbounded agent loops.

Haystack component or direct REST API

Does your pipeline already return Document objects? Add these components directly. They normalize tweet text, metadata, URLs, and pagination.

Use direct TwexAPI REST routes or the Python SDK for followers, following, replies, quotes, lists, communities, media, trends, and approved writes.

Both approaches use the same TwexAPI contracts. Components only support tweet search and user timelines.

MCP handoff alternative

When an agent discovers endpoints first, use Twexapi MCP for collection and convert rows into Haystack Document objects manually. Prefer the Haystack components when the pipeline already runs inside Haystack and you want typed pagination without agent discovery.

Source and contracts

Was this page helpful?