Regex Evaluation – Arxiv ID

You can extract a pattern as an additional field with Pentaho using the “Regex evaluation” step.

Example to extract the Arxiv-ID with this Regex: .*(\d{4}\.\d{4,5}).*

The found regex will be added as new field to the stream.

Advertisement

Leave a Reply

Fill in your details below or click an icon to log in:

WordPress.com Logo

You are commenting using your WordPress.com account. Log Out /  Change )

Twitter picture

You are commenting using your Twitter account. Log Out /  Change )

Facebook photo

You are commenting using your Facebook account. Log Out /  Change )

Connecting to %s