relational database
This commit is contained in:
@@ -82,6 +82,40 @@ and repository name as key and add that dictionary to db. For example if I am ad
|
|||||||
packages db, I am actually adding, a Dictionary with ["repotype", foo] to db.
|
packages db, I am actually adding, a Dictionary with ["repotype", foo] to db.
|
||||||
This is cheap, bloated and slow. We must have different repository databases for packages.
|
This is cheap, bloated and slow. We must have different repository databases for packages.
|
||||||
|
|
||||||
|
* We should really consider using a simple and fast relational database (sqlite) for database operations. Here
|
||||||
|
is the rationale and some issues.
|
||||||
|
- Current implementation uses berkeleydb as a storage of dictionaries, with keys and Pickled Python objects.
|
||||||
|
This approach simplifies coding requirements but leads to some problems.
|
||||||
|
1. Some operations cannot be done effectively. Searching can only be done by iterating all objects and searching
|
||||||
|
one by one. component, reverse dependencies etc are holding lists of package names that are updated on every
|
||||||
|
package operation, Multiple repository problem is
|
||||||
|
|
||||||
|
2.Performance: although fast, pickling and unpickling whole objects is still a time consuming process.
|
||||||
|
Pickled Objects contains object metadata and are larger then the data they hold, so reading, writing pickled
|
||||||
|
Objects to db is slower than reading and writing only necessary fields. large objects (Order of kilobytes)
|
||||||
|
causes berkeleydb data buckets to overflow and caches works less efficently.
|
||||||
|
|
||||||
|
4.Disk usage: Using pickled objects requires more memory and disk space. Total size of a current pisi database
|
||||||
|
is over 100Mb, if we modify files.bdb to only hold filenames and packages, size is still around 50MB.
|
||||||
|
|
||||||
|
Rational Databases offers number of advantages and disadvantages:
|
||||||
|
|
||||||
|
Advantages:
|
||||||
|
+ Relational databases offers much more flexible operations on stored data. Searching, grouping, sorting
|
||||||
|
can be done by database engine efficently.
|
||||||
|
+ There will be only a single database file and much smaller database size compared to berkeleydb files
|
||||||
|
holding pickled python objects.
|
||||||
|
+ Since database no longer holds pickled python objects, it is possible to read and modify the database
|
||||||
|
by other programs written in different languages (Java, C++ etc.)
|
||||||
|
+ SQL lite is generally very fast, probably beating "Berkeleydb with pickled objects" by the order of 4-5 on
|
||||||
|
queries.
|
||||||
|
|
||||||
|
Disadvantages
|
||||||
|
- Using a relational database will require more coding when adding and retrieving contents of objects.
|
||||||
|
- Introducing SQL language makes code more complex and harder to maintain.
|
||||||
|
- Designing a good schema, identifying required indexes etc. takes time
|
||||||
|
- Some developers are superstitious about sql and relational databases ;)
|
||||||
|
|
||||||
==> class attributes
|
==> class attributes
|
||||||
* In some classes there are some attributes assigned but never
|
* In some classes there are some attributes assigned but never
|
||||||
used. (see remove_unused_attributes.patch)
|
used. (see remove_unused_attributes.patch)
|
||||||
|
|||||||
Reference in New Issue
Block a user