Languages

Parsing JSON and XML data with Python

JGJaya Gupta21 Mar 2023 Β· Updated 04 Oct 2026 Β· 8 min read
Parsing JSON and XML data with Python

Quick answer: Python parses JSON with the built-in json module (json.loads() for strings, json.load() for files) and XML with xml.etree.ElementTree (ET.fromstring() and ET.parse()). Both modules ship with Python 3, so you need no extra installation. JSON becomes ordinary dictionaries and lists; XML becomes a tree of Element objects you walk with find(), findall() and iter().

JSON (JavaScript Object Notation) and XML (eXtensible Markup Language) are the two formats you will meet most often when a script talks to a REST API, reads a configuration file, or pulls data from a network device. In this guide you will learn how to parse both with Python’s standard library, how to write data back out, when to choose one over the other, and the mistakes that trip up beginners. The examples use network-device data because that is where this topic comes up constantly, but the techniques apply to any JSON or XML you encounter.

Why JSON and XML matter for Python developers

Almost every modern API β€” cloud platforms, GitHub, payment gateways, network controllers such as Cisco DNA Center β€” returns JSON. It is compact, maps directly onto Python’s dict and list types, and is easy to read.

XML is older and more verbose, but it is far from dead. Network protocols like NETCONF, many enterprise SOAP services, Android layouts, Maven pom.xml files, RSS feeds and countless legacy configuration systems still speak XML. A working engineer needs to be comfortable with both.

If you are new to the language itself, start with What is Python programming? and come back here; everything below assumes you know basic dictionaries and loops.

Parsing JSON with the json module

The json module has four functions you will use daily:

  • json.loads(string) – parse a JSON string into Python objects.
  • json.load(file_object) – parse JSON from an open file.
  • json.dumps(obj) – convert Python objects to a JSON string.
  • json.dump(obj, file_object) – write Python objects to a file as JSON.

Here is the device example from the original post, parsed and inspected:

import json

json_data = '''
{
  "name": "Switch1",
  "status": "OK",
  "ports": [
    {"name": "Port1", "status": "UP",   "speed": "1G"},
    {"name": "Port2", "status": "DOWN", "speed": "1G"}
  ]
}
'''

# Parse the JSON string into a Python dict
data = json.loads(json_data)

print("Device name:", data["name"])
print("Device status:", data["status"])
print("Number of ports:", len(data["ports"]))

for port in data["ports"]:
    print(f'{port["name"]}: {port["status"]} ({port["speed"]})')

Output:

Device name: Switch1
Device status: OK
Number of ports: 2
Port1: UP (1G)
Port2: DOWN (1G)

Notice there is nothing special about data after parsing β€” it is a plain dictionary. JSON objects become dict, arrays become list, strings stay strings, numbers become int or float, true/false become True/False, and null becomes None.

Reading from and writing to a file

import json

# Read
with open("inventory.json", encoding="utf-8") as f:
    inventory = json.load(f)

# Modify: mark every DOWN port for review
for port in inventory["ports"]:
    if port["status"] == "DOWN":
        port["needs_review"] = True

# Write back, pretty-printed with 2-space indent
with open("inventory_reviewed.json", "w", encoding="utf-8") as f:
    json.dump(inventory, f, indent=2, sort_keys=True)

The indent argument is what people mean by “pretty-printing”. Without it, json.dump() writes everything on one line, which is fine for machines but painful for humans.

Parsing XML with xml.etree.ElementTree

XML data is a tree: one root element containing child elements, each of which can hold text, attributes and more children. ElementTree (imported as ET by convention) gives you that tree directly.

import xml.etree.ElementTree as ET

xml_data = '''
<device>
  <name>Switch1</name>
  <status>OK</status>
  <ports>
    <port id="1">
      <name>Port1</name>
      <status>UP</status>
      <speed>1G</speed>
    </port>
    <port id="2">
      <name>Port2</name>
      <status>DOWN</status>
      <speed>1G</speed>
    </port>
  </ports>
</device>
'''

root = ET.fromstring(xml_data)        # root is the <device> element

print("Device name:", root.find("name").text)
print("Device status:", root.find("status").text)

ports = root.findall(".//port")       # XPath-style: every <port> anywhere below root
print("Number of ports:", len(ports))

for port in ports:
    print(f'Port {port.get("id")}: {port.find("name").text} is {port.find("status").text}')

Key methods to remember:

  • find(path) returns the first matching child element (or None).
  • findall(path) returns a list of all matches. The .// prefix searches at any depth.
  • .text gives the text inside an element; .get("attr") reads an attribute such as id="1".
  • ET.parse("file.xml").getroot() does the same job for a file on disk.

To create or modify XML, use ET.SubElement(parent, "tag"), set .text, then call ET.ElementTree(root).write("out.xml", encoding="utf-8", xml_declaration=True). Python 3.9+ also has ET.indent(tree) for pretty-printing before writing.

JSON vs XML: which should you use?

Aspect JSON XML
Python module json xml.etree.ElementTree
Parsed into dict, list, str, int, float, bool, None Tree of Element objects
Verbosity Low High (opening and closing tags)
Attributes Not supported (keys only) Supported (<port id="1">)
Comments Not allowed Allowed
Schema validation JSON Schema (third-party) XSD, DTD (mature tooling)
Typical use REST APIs, config files, logs NETCONF, SOAP, legacy enterprise systems, documents

Rule of thumb: pick JSON for anything new you control. Use XML when the system on the other end demands it. In network automation specifically, REST-based controllers return JSON while NETCONF/YANG-based devices return XML, so most engineers end up using both in the same project. Our guide on why Python is the language of choice for network engineering explains where each fits in a real workflow.

Handling real-world messiness

Sample data in tutorials is always clean. Production data is not. Three habits save hours of debugging:

Catch parse errors. Both modules raise exceptions on malformed input β€” json.JSONDecodeError and ET.ParseError. Wrap parsing in try/except when the data comes from outside your program.

Use .get() for optional keys. data.get("uptime", 0) returns a default instead of raising KeyError when a field is missing.

Watch for XML namespaces. Real NETCONF or SOAP responses often look like <ns0:device xmlns:ns0="urn:example">. In that case root.find("name") returns None; you must search with the namespace: root.find("{urn:example}name"), or pass a namespace map to find().

Common mistakes when parsing JSON and XML

  1. Mixing up load and loads. The s means string. Passing a file object to loads() or a string to load() raises a confusing TypeError.
  2. Treating JSON as Python. JSON requires double quotes, has no trailing commas and no comments. A file that looks almost right will still fail to parse.
  3. Forgetting .text on XML elements. root.find("name") returns an Element, not a string. Printing it gives <Element 'name' at 0x...>.
  4. Using find() when you need findall(). find() silently returns only the first match, so loops over it process one item.
  5. Parsing untrusted XML with the default parser. Maliciously crafted XML can trigger “billion laughs” style attacks. For untrusted input use the defusedxml package, which wraps ElementTree safely.
  6. Ignoring encoding. Always open files with encoding="utf-8"; the default varies by operating system and breaks on non-ASCII characters.

Frequently asked questions

Do I need to install anything to parse JSON or XML in Python?

No. Both json and xml.etree.ElementTree are part of the Python 3 standard library. Third-party options such as lxml (faster, full XPath) or orjson (faster JSON) are optional upgrades for heavy workloads.

How do I convert XML to JSON in Python?

There is no built-in one-liner because XML attributes and mixed content have no direct JSON equivalent. The common approach is to walk the tree with ElementTree, build a dictionary by hand, then call json.dumps(). The third-party xmltodict package automates this for most documents.

What is the difference between json.loads() and requests.get().json()?

The requests library’s .json() method simply calls json.loads() on the response body for you. Use it when you fetch data from an API; use json.loads() when you already have the string.

Can ElementTree handle very large XML files?

ET.parse() loads the whole document into memory. For multi-gigabyte files use ET.iterparse(), which streams elements one at a time and lets you call elem.clear() to free memory as you go.

Key takeaways

  • json.loads()/json.load() turn JSON into plain Python dictionaries and lists; json.dumps()/json.dump() go the other way.
  • ET.fromstring()/ET.parse() turn XML into an element tree; use find(), findall(".//tag"), .text and .get() to read it.
  • Prefer JSON for new APIs and configs; expect XML from NETCONF, SOAP and legacy systems.
  • Handle parse errors, missing keys and XML namespaces explicitly β€” production data is never as clean as the examples.
  • Use defusedxml for untrusted XML and iterparse() for very large files.

Want to go from parsing a sample file to building complete applications that consume and expose APIs? Our Full Stack Development course covers Python, REST APIs, databases and deployment with live projects and placement support. Prefer video? Follow along on our YouTube channel.

JG
Written byJaya Gupta

Part of the Techknowledgehub team of industry mentors, writing practical guides to help you build a job-ready tech career.

More articles by Jaya Gupta β†’
Keep reading

Related articles

Leave a Reply