Research Data

The Research Data module allows you to build networks from scholarly research literature — for example, a citation network around a set of key papers, or a co-authorship network around a group of researchers. You configure a Data Pull describing what to collect, run it, and Polinode collects the data and produces a network for you, ready to explore like any other network in Polinode.

Please note that the Research Data module needs to be enabled for your organization before you will be able to use it — it is not available on individual accounts. If you would like it enabled please contact your Account Manager or support@polinode.com. Once enabled, a Research Data item will appear in the left-hand side menu. Collecting research data consumes Data Credits — see Data Credits and Costs below.

The Research Data Screen

Clicking on Research Data in the left-hand side menu shows a list of your saved Data Pulls. Towards the top of the screen you will see your current Data Credit balance along with an Add Data Credits button, a "Search data pulls" input and a New Data Pull button. The list has the following columns:

  1. Name: The name of the Data Pull. Hovering over the name shows a summary of its seeds, e.g. "3 paper seeds".
  2. Source: The data source — currently Research literature.
  3. Graph: The type of network the pull produces — Paper citations or Co-authorship.
  4. Last run: The status of the most recent run — queued, running, succeeded, succeeded (capped) or failed — or "Never run".
  5. Modified: The date the Data Pull was last modified.
  6. Access: Your permission level on the Data Pull — Owner, Edit or View.

Each row also has action buttons to Run the pull, view its Run history, Edit it, Manage users and Delete it. Buttons you do not have the necessary access for are disabled — hovering over one explains why. While a run is in progress the Run, Edit and Delete buttons are disabled until it finishes. Deleting a Data Pull removes it and its run history but keeps any networks it has already produced.

Visibility and Sharing

Data Pulls are private to you by default. Enabling the Research Data module for your organization does not make everyone's Data Pulls visible to everyone else: you see only the Data Pulls you created or that somebody has shared with you, and the same applies to their run history and Excel exports. This matches how Networks and Surveys work.

To share a Data Pull, click the Manage users button on its row and then Add User. You can grant one of three permission levels:

  1. View: Can open the Data Pull, see its run history and download run exports, but cannot change or run it.
  2. Edit: Everything in View, plus editing the Data Pull's configuration and running it. Because every run spends your organization's shared Data Credits, running requires Edit rather than View.
  3. Owner: The creator of the Data Pull. Only the Owner can delete it. Ownership cannot be granted through sharing.

You can also give a user the ability to add and remove other users. As with Networks, you cannot grant somebody a level of access higher than your own, and the Owner's access cannot be edited or removed. Users must be members of the same organization as the Data Pull, and they receive an email letting them know it has been shared with them.

Note that Data Credits, the limit on how many Data Pulls can run at the same time, and whether the module is available at all remain organization-wide.

Creating a Data Pull

Click on the New Data Pull button and a drawer will open on the right-hand side of the screen. Give the Data Pull a name and then work through the following sections.

Data Source

  1. Graph type: Choose between Paper citations — a network where the nodes are papers and a directed edge links each citing paper to the paper it references — and Co-authorship — a network where the nodes are authors and an edge links authors who have published together, weighted by the number of shared papers.
  2. Seed entity: Choose whether to start from Papers (identified by DOI or OpenAlex work ID) or Authors (identified by ORCID or OpenAlex author ID). Any combination of graph type and seed entity is allowed — for example, you can seed a citation network from a set of authors, in which case the authors' publications become the starting papers.

Seed Identifiers

The seeds are the papers or authors the collection starts from. You can supply them in one of two ways:

  1. Upload a file: Click Upload Excel to upload a spreadsheet of identifiers. A Download template link provides the expected format. The identifier column accepts either scheme for the seed entity you have chosen, and the two can be mixed freely within the same column — OpenAlex IDs may be given either bare (W2741809807) or as a full https://openalex.org/... URL. After uploading you will see how many valid identifiers were found, along with any rows that could not be parsed.
  2. From an existing network: Select one of your existing networks and the attribute column that contains the identifiers (DOIs or OpenAlex work IDs for papers, ORCIDs or OpenAlex author IDs for authors). The column's values are read and validated when you save.

Publication Filters

You can optionally restrict which papers are included using a Publication start date, a Publication end date and a Minimum citations threshold (only papers with at least this many citations are included). These filters apply to the papers collected at every step, not just the seeds.

Node Attributes

For a citation graph you can also choose which node attributes each paper carries. The standard attributes are always included: DOI, publication details, citation counts, authors, source and primary topic, together with FWCI (field-weighted citation impact), Citation Percentile, Download Link (a link to the open-access full text, where one exists) and Publisher Page. The text-heavy fields — Abstract and Keywords — are optional and off by default; tick them to include them as node attributes.

Expansion

Expansion controls how far the collection spreads out from your seeds:

  1. For a citation graph, Citing paper hops and Referenced paper hops (each 0–3) control how many hops to expand through papers that cite, and papers that are referenced by, the current set of papers.
  2. For a co-authorship graph, Co-author hops (0–3) controls how many hops to expand through authors who have published with the current set of authors.
  3. Min. connections: A discovered paper or author is only added to the network when it connects to at least this many nodes in the current set. Raising this keeps the network focused on well-connected material.
  4. Max nodes per hop: A hard cap on how many new nodes can be admitted at each hop. When more candidates qualify than the cap allows, the best-connected and most-cited candidates are kept.

Usage Limits

Max Data Credits per run is a hard stop: the run ends once it has used this many Data Credits, and whatever was reserved but not spent is refunded. New Data Pulls default to a limit of 1,000 credits. You can clear the limit entirely, in which case a run will reserve your full available balance as its spending limit instead. If a run hits its limit it completes with the network collected so far and is marked "succeeded (capped)".

Once you are done, click Create (or Save) to save the Data Pull, or Create and Run to save it and start a run immediately.

Running a Data Pull

When you run a Data Pull, a confirmation dialog summarizes what will happen: the run reserves Data Credits up front — the pull's credit limit if one is set, otherwise your full available balance — and refunds whatever it doesn't spend when it finishes. If you don't have enough credits available to cover the reservation the dialog will tell you how many more you need.

The dialog also has a Use cached data where available switch, which is on by default. Where Polinode has recently collected the same records, they are reused at a discount and the savings are reported on the run. Turn the switch off if you want fresh data only, with every record collected again at the standard price.

Runs happen in the background and can take a while for larger pulls — you will receive an email when your network is ready. Each Data Pull can only have one run in progress at a time, and an organization can have at most two Research Data runs in progress at once.

Run History

The Run history button opens a timeline of the pull's most recent runs (up to 50). Each run shows its status, when it ran, progress (node and edge counts), the Data Credits spent against the amount reserved, any savings from cached data, and any error if the run failed. For a successful run you can jump straight to the produced network via the Open network link, or download the run as an Excel workbook using the download button.

The Excel download contains three sheets: About This Data, which records exactly what was collected and how (the pull's configuration, filters, seeds, whether the run was capped and the credits charged, together with source attribution); Nodes; and Edges.

The Produced Network

Each run produces a new Private network named after the Data Pull and the run date. For a paper citation network, nodes are papers labelled by title, with attributes including DOI, Publication Date, Publication Year, Work Type, Language, Citation Count, Citation Percentile, FWCI, Source (the journal or venue), Open Access, Download Link, Publisher Page, Retracted, Author Count, Authors, Primary Topic, whether the paper was a Seed and its Citation Hop — plus the Abstract and Keywords if you selected them (see Node Attributes above). Edges are directed from the citing paper to the cited paper.

For a co-authorship network, nodes are authors labelled by name, with attributes including ORCID, Works Count, Citation Count, H-index, Institution, Institution Country, whether the author was a Seed and their Author Hop. There is one edge per pair of co-authors, weighted by their number of shared papers, with First Collaboration and Latest Collaboration date attributes.

The usual platform limits of 50,000 nodes and 250,000 edges per network apply.

Data Credits and Costs

Research Data shares the same prepaid Data Credits balance as the Social Data module. A run is charged 0.01 credits per unique node added to the network (a paper or an author), and edges are free. Credits are priced at $0.01 each, so 1 credit builds 100 nodes and $1 builds 10,000 nodes. A typical citation or co-authorship network of a few thousand nodes therefore costs well under a dollar. Records served from Polinode's recent-collection cache are charged at half the standard rate, and the savings are shown on the run and in your completion email. Because credits are reserved up front and unspent credits are refunded when the run settles, the reservation you see when starting a run is an upper bound, not the expected cost. See What your Data Credits get you for the full pricing summary.

To add credits, click Add Data Credits on the Research Data screen. Admin users in your organization can have credits added immediately with an invoice to follow, while other members' requests are sent to your Polinode account team. Data Credits never expire.