MsPASS C++ API  2.4.4.dev114+gf4c3cfaca
Defines the C++ API for MsPASS
Loading...
Searching...
No Matches
Public Member Functions | Public Attributes | Protected Attributes | List of all members
mspass::utility::ProcessingHistory Class Reference

Lightweight class to preserve procesing chain of atomic objects. More...

#include <ProcessingHistory.h>

Inheritance diagram for mspass::utility::ProcessingHistory:
Inheritance graph
[legend]
Collaboration diagram for mspass::utility::ProcessingHistory:
Collaboration graph
[legend]

Public Member Functions

 ProcessingHistory ()
 
 ProcessingHistory (const std::string jobnm, const std::string jid)
 
 ProcessingHistory (const std::string jobnm, const std::string jid, const std::multimap< std::string, NodeData > &nodesin, const NodeData &current, const ErrorLogger &elogin)
 
 ProcessingHistory (const ProcessingHistory &parent)
 
bool is_empty () const
 
bool is_raw () const
 
bool is_origin () const
 
bool is_volatile () const
 
bool is_saved () const
 
size_t number_of_stages () override
 Return number of processing stages that have been applied to this object.
 
void set_as_origin (const std::string alg, const std::string algid, const std::string uuid, const AtomicType typ, bool define_as_raw=false)
 
std::string new_ensemble_process (const std::string alg, const std::string algid, const AtomicType typ, const std::vector< ProcessingHistory * > parents, const bool create_newid=true)
 
void add_one_input (const ProcessingHistory &data_to_add)
 Add one datum as an input for current data.
 
void add_many_inputs (const std::vector< ProcessingHistory * > &d)
 Define several data objects as inputs.
 
void merge (const ProcessingHistory &data_to_add)
 Merge the history nodes from another.
 
void accumulate (const std::string alg, const std::string algid, const AtomicType typ, const ProcessingHistory &newinput)
 Method to use with a spark reduce algorithm.
 
std::string clean_accumulate_uuids ()
 Clean up inconsistent uuids that can be produced by reduce.
 
std::string new_map (const std::string alg, const std::string algid, const AtomicType typ, const ProcessingStatus newstatus=ProcessingStatus::VOLATILE)
 Define this algorithm as a one-to-one map of same type data.
 
std::string new_map (const std::string alg, const std::string algid, const AtomicType typ, const ProcessingHistory &data_to_clone, const ProcessingStatus newstatus=ProcessingStatus::VOLATILE)
 Define this algorithm as a one-to-one map.
 
std::string map_as_saved (const std::string alg, const std::string algid, const AtomicType typ)
 Prepare the current data for saving.
 
void clear ()
 
std::multimap< std::string, mspass::utility::NodeData > get_nodes () const
 
int stage () const
 
ProcessingStatus status () const
 
std::string id () const
 
std::pair< std::string, std::string > created_by () const
 
NodeData current_nodedata () const
 
std::string newid ()
 
int number_inputs () const
 
int number_inputs (const std::string uuidstr) const
 
void set_id (const std::string newid)
 
std::list< mspass::utility::NodeData > inputs (const std::string id_to_find) const
 Return a list of data that define the inputs to a give uuids.
 
ProcessingHistory & operator= (const ProcessingHistory &parent)
 
- Public Member Functions inherited from mspass::utility::BasicProcessingHistory
 BasicProcessingHistory (const std::string jobname, const std::string jobid)
 
 BasicProcessingHistory (const BasicProcessingHistory &parent)
 
std::string jobid () const
 
void set_jobid (const std::string &newjid)
 
std::string jobname () const
 
void set_jobname (const std::string jobname)
 
BasicProcessingHistory & operator= (const BasicProcessingHistory &parent)
 

Public Attributes

ErrorLogger elog
 

Protected Attributes

std::multimap< std::string, mspass::utility::NodeData > nodes
 
- Protected Attributes inherited from mspass::utility::BasicProcessingHistory
std::string jid
 
std::string jnm
 

Detailed Description

Lightweight class to preserve procesing chain of atomic objects.

This class is intended to be used as a parent for any data object in MsPASS that should be considered atomic. It is designed to completely preserve the chain of processing algorithms applied to any atomic data to put it in it's current state. It is designed to save that information during processing with the core information that can then be saved to define the state. Writers for atomic objects inheriting this class should arrange to save the data contained in it to history collection in MongoDB. Note that actually doing the inverse is a different problem that are expected to be implemented as extesions of this class to be used in special programs used to reconstrut a data workflow and the processing chain applied to produce any final output.

The design was complicated by the need to keep the history data from causing memory bloat. A careless implementation could be prone to that problem even for modest chains, but we were particularly worried about iterative algorithms that could conceivably multiply the size of out of control. There was also the fundamental problem of dealing with transient versus data stored in longer term storage instead of just in memory. Our implementation was simplified by using the concept of a unique id with a Universal Unique IDentifier. (UUID) Our history mechanism assumes each data object has a uuid assigned to it on creation by an implementation id of the one object this particular record is associated with on dependent mechanism. That is, whenever a new object is created in MsPASS using the history feature one of these records will be created for each data object that is defined as atomic. This string defines unique key for the object it could be connected to with the this pointer. The parents of the current object are defined by the inputs data structure below.

In the current implementation id is string representation of a uuid maintained by each atomic object. We use a string to maximize flexibility at a minor cost for storage.

Names used imply the following concepts: raw - means the data is new input to mspass (raw data from data center, field experiment, or simulation). That tag means no prior history can be reconstructed. origin - top-level ancestor of current data. The top of a processing chain is always tagged as an origin. A top level can also be "raw" but not necessarily. In particular, readers that load partially processed data should mark the data read as an origin, but not raw. stage - all processed data objects that are volatile elements within a workflow are defined as a stage. They are presumed to leave their existence known only through ancestory preserved in the processing chain. A stage becomes a potential root only when it is saved by a writer where the writer will mark that position as a save. Considered calling this a branch, but that doesn't capture the concept right since we require this mechanism to correctly perserve splits into multiple outputs. We preserve that cleanly for each data object. That is, the implementation make it easy to reconstruct the history of a single final data object, but reconstructing interlinks between objects in an overall processing flow will be a challenge. That was a necessary compomise to avoid memory bloat. The history is properly viewed as a tree branching from a single root (the final output) to leaves that define all it's parents.

The concepts of raw, origin, and stage are implemented with the enum class defined above called ProcessingStatus. Each history record has that as an attribute, but each call to new_stage updates a copy kept inside this object to simplify the python wrappers.

Constructor & Destructor Documentation

◆ ProcessingHistory() [1/4]

mspass::utility::ProcessingHistory::ProcessingHistory ( )

Default constructor.

82 : elog() {
83 current_status = ProcessingStatus::UNDEFINED;
84 current_id = "UNDEFINED";
85 current_stage =
86 -1; // illegal value that could be used as signal for uninitalized
87 mytype = AtomicType::UNDEFINED;
88 algorithm = "UNDEFINED";
89 algid = "UNDEFINED";
90}
ErrorLogger elog
Definition ProcessingHistory.h:246

◆ ProcessingHistory() [2/4]

mspass::utility::ProcessingHistory::ProcessingHistory ( const std::string  jobnm,
const std::string  jid 
)

Construct and fill in BasicProcessingHistory job attributes.

Parameters
jobnm- set as jobname
jid- set as jobid
93 : BasicProcessingHistory(jobnm, jid), elog() {
94 current_status = ProcessingStatus::UNDEFINED;
95 current_id = "UNDEFINED";
96 current_stage =
97 -1; // illegal value that could be used as signal for uninitalized
98 mytype = AtomicType::UNDEFINED;
99 algorithm = "UNDEFINED";
100 algid = "UNDEFINED";
101}
std::string jid
Definition ProcessingHistory.h:105

◆ ProcessingHistory() [3/4]

mspass::utility::ProcessingHistory::ProcessingHistory ( const std::string  jobnm,
const std::string  jid,
const std::multimap< std::string, NodeData > &  nodesin,
const NodeData &  current,
const ErrorLogger &  elogin 
)

Restore all state from a field-level persistence representation.

106 : BasicProcessingHistory(jobnm, jid), elog(elogin), nodes(nodesin),
107 algorithm(current.algorithm), algid(current.algid) {
108 current_status = current.status;
109 current_id = current.uuid;
110 current_stage = current.stage;
111 mytype = current.type;
112}
std::multimap< std::string, mspass::utility::NodeData > nodes
Definition ProcessingHistory.h:676

References mspass::utility::NodeData::stage, mspass::utility::NodeData::status, mspass::utility::NodeData::type, and mspass::utility::NodeData::uuid.

◆ ProcessingHistory() [4/4]

mspass::utility::ProcessingHistory::ProcessingHistory ( const ProcessingHistory &  parent)

Standard copy constructor.

114 : BasicProcessingHistory(parent), elog(parent.elog), nodes(parent.nodes),
115 algorithm(parent.algorithm), algid(parent.algid) {
116 current_status = parent.current_status;
117 current_id = parent.current_id;
118 current_stage = parent.current_stage;
119 mytype = parent.mytype;
120}

Member Function Documentation

◆ accumulate()

void mspass::utility::ProcessingHistory::accumulate ( const std::string  alg,
const std::string  algid,
const AtomicType  typ,
const ProcessingHistory &  newinput 
)

Method to use with a spark reduce algorithm.

A reduce operator in spark utilizes a binary function where two inputs are used to generate a single output object. Because the inputs could be scattered on multiple processor nodes this operation must be associative. The new_ensemble_process method does not satisfy that constraint so this method was necessary to handle that type of algorithm correctly.

The way this algorithm works is it fundamentally branches on two different cases: (1) initialization, which is detected by testing if the node data map is empty or (2) secondary calls. This should work even if multiple inputs are combined at the end of the reduce operation because the copies being merged will not be empty. Note an empty input will create a complaint entry in the error log.

489 {
490 ProcessingHistory newinput(ni);
491 if ((newinput.algorithm != algin) || (newinput.algid != algidin) ||
492 (newinput.jid != newinput.jobid()) ||
493 (newinput.jnm != newinput.jobname())) {
494 NodeData nd;
495 nd = newinput.current_nodedata();
496 newinput.newid();
497 pair<string, NodeData> pn(newinput.current_id, nd);
498 newinput.nodes.insert(pn);
499 newinput.jid = newinput.jobid();
500 newinput.jnm = newinput.jobname();
501 newinput.algorithm = algin;
502 newinput.algid = algidin;
503 newinput.current_status = ProcessingStatus::VOLATILE;
504 newinput.current_stage = nd.stage + 1;
505 newinput.mytype = typ;
506 }
507 /* We have to detect an initialization condition without losing the
508 stored history. There are two conditions we need to handle. First,
509 if we create an empty container to hold the accmulator and put it on the
510 left hand side we will want to clear the history chain or we will
511 accumulate random junk. The second condition is if we accumulate in
512 a way were the left hand side is some existing data where we do want to
513 preserve the history. For the is_empty logic: we just copy the
514 newinput's history and add make its current node data the connection
515 backward - i.e. we have to make a new uuid and add an entry. */
516 if (this->is_empty()) {
517 this->newid();
518 nodes = ni.get_nodes();
519 NodeData nd;
520 nd = ni.current_nodedata();
521 pair<string, NodeData> pn(current_id, nd);
522 this->nodes.insert(pn);
523 this->set_jobid(ni.jobid());
524 this->set_jobname(ni.jobname());
525 algorithm = algin;
526 algid = algidin;
527 current_status = ProcessingStatus::VOLATILE;
528 current_stage = nd.stage + 1;
529 mytype = typ;
530 }
531 /* This is the condition for a left hand side that is not empty but not
532 yet initialized. We detect this condition by a mismatch in all the unique
533 names and ids that mark the current process define this reduce operation*/
534 else if ((this->algorithm != algin) || (this->algid != algidin) ||
535 (this->jid != newinput.jobid()) ||
536 (this->jnm != newinput.jobname())) {
537 /* This is similar to the block above, but the key difference here is we
538 have to push this's history data to convert it's current data to define an
539 input. That means getting a new uuid and pushing current node data to the
540 nodes map as an input */
541 NodeData nd;
542 nd = this->current_nodedata();
543 this->newid();
544 pair<string, NodeData> pn(current_id, nd);
545 this->nodes.insert(pn);
546 this->jid = newinput.jobid();
547 this->jnm = newinput.jobname();
548 this->algorithm = algin;
549 this->algid = algidin;
550 this->current_status = ProcessingStatus::VOLATILE;
551 this->current_stage = nd.stage + 1;
552 this->mytype = typ;
553 this->merge(newinput);
554 } else {
555 this->merge(newinput);
556 }
557}
std::string jnm
Definition ProcessingHistory.h:107
void set_jobid(const std::string &newjid)
Definition ProcessingHistory.h:89
void set_jobname(const std::string jobname)
Definition ProcessingHistory.h:93
NodeData current_nodedata() const
Definition ProcessingHistory.cc:672
void merge(const ProcessingHistory &data_to_add)
Merge the history nodes from another.
Definition ProcessingHistory.cc:452
ProcessingHistory()
Definition ProcessingHistory.cc:82
bool is_empty() const
Definition ProcessingHistory.cc:121
std::string newid()
Definition ProcessingHistory.cc:664

References current_nodedata(), get_nodes(), is_empty(), mspass::utility::BasicProcessingHistory::jid, mspass::utility::BasicProcessingHistory::jnm, mspass::utility::BasicProcessingHistory::jobid(), mspass::utility::BasicProcessingHistory::jobname(), merge(), newid(), nodes, mspass::utility::BasicProcessingHistory::set_jobid(), mspass::utility::BasicProcessingHistory::set_jobname(), and mspass::utility::NodeData::stage.

◆ add_many_inputs()

void mspass::utility::ProcessingHistory::add_many_inputs ( const std::vector< ProcessingHistory * > &  d)

Define several data objects as inputs.

This method acts like add_one_input in that it alters only the inputs chain. In fact it is nothing more than a loop over the components of the vector calling add_one_input for each component.

Parameters
dis the vector of data to define as inputs
316 {
317 vector<ProcessingHistory *>::const_iterator dptr;
318 for (dptr = d.begin(); dptr != d.end(); ++dptr) {
320 ptr = (*dptr);
321 this->add_one_input(*ptr);
322 }
323}
void add_one_input(const ProcessingHistory &data_to_add)
Add one datum as an input for current data.
Definition ProcessingHistory.cc:275

References add_one_input().

◆ add_one_input()

void mspass::utility::ProcessingHistory::add_one_input ( const ProcessingHistory &  data_to_add)

Add one datum as an input for current data.

This method MUST ONLY be called after a call to new_ensemble_process in the situation were additional inputs need to be defined that were not available at the time new_ensemble_process was called. An example might be a stack that was created within the scope of "algorithm" and then used in some way to create the output data. In any case it differs fundamentally from new_ensemble_process in that it does not touch attributes that define the current state of "this". It simply says this is another input to the data "this" contains.

Parameters
data_to_addis the ProcessingHistory of the data object to be defined as input. Note the type of the data to which it is linked will be saved as the base of the input chain from data_to_add. It can be different from the type of "this".
275 {
276
277 if (data_to_add.is_empty()) {
278 stringstream ss;
279 ss << "Data with uuid=" << data_to_add.id() << " has an empty history chain"
280 << endl
281 << "At best this will leave ProcessingHistory incomplete" << endl;
282 elog.log_error("ProcessingHistory::add_one_input", ss.str(),
283 ErrorSeverity::Complaint);
284 } else {
285 multimap<string, NodeData>::iterator nptr;
286 multimap<string, NodeData> newhistory = data_to_add.get_nodes();
287 multimap<string, NodeData>::iterator nl, nu;
288 /* As above this one needs check for duplicates and only add
289 a node if the data are unique. This is simple compared to
290 new_ensemble_process because we just have to check one object's history at a
291 time. */
292 for (nptr = newhistory.begin(); nptr != newhistory.end(); ++nptr) {
293 string key(nptr->first);
294 if (this->nodes.count(key) > 0) {
295 nl = this->nodes.lower_bound(key);
296 nu = this->nodes.upper_bound(key);
297 for (auto ptr = nl; ptr != nu; ++ptr) {
298 NodeData ndtest(ptr->second);
299 if (ndtest != (nptr->second)) {
300 this->nodes.insert(*nptr);
301 }
302 }
303 } else {
304 this->nodes.insert(*nptr);
305 }
306 }
307 /* Don't forget head node data*/
308 NodeData nd = data_to_add.current_nodedata();
309 NodeData ndhere = this->current_nodedata();
310 pair<string, NodeData> pnd(current_id, nd);
311 this->nodes.insert(pnd);
312 }
313}
int log_error(const mspass::utility::MsPASSError &merr)
Definition ErrorLogger.cc:72

References current_nodedata(), elog, get_nodes(), id(), is_empty(), mspass::utility::ErrorLogger::log_error(), and nodes.

◆ clean_accumulate_uuids()

string mspass::utility::ProcessingHistory::clean_accumulate_uuids ( )

Clean up inconsistent uuids that can be produced by reduce.

In a spark reduce operation it is possible to create multiple uuid keys for inputs to the same algorithm instance. That happpens because the mechanism used by ProcessingHistory to define the process history tree is not associative. When a reduce gets sprayed across multiple nodes multiple initializations can occur that make artifical inconsitent uuids. This method should normally be called after a reduce operator if history is being preserved or the history chain may be foobarred - no invalid just mess up with extra branches in the processing tree.

A VERY IMPORTANT limitation of the algorithm used by this method is that the combination of algorithm and algid in "this" MUST be unique for a given job run when a reduce is called. i.e. if an earlier workflow had used alg and algid but with a different jobid and jobname the distintion cannot be detected with this algorithm. This means our global history handling must guarantee algid is unique for each run.

Returns
unique uuid for alg,algid match set in the history chain. Note if there are no duplicates it simply returns the only one it finds. If there are duplicates it returns the lexically smallest (first in alphabetic order) uuid. Most importantly if there is no match or if history is empty it returns the string UNDEFINED.
559 {
560 /* Return undefined immediately if the history chain is empty */
561 if (this->is_empty())
562 return string("UNDEFINED");
563 NodeData ndthis = this->current_nodedata();
564 string alg(ndthis.algorithm);
565 string algidtest(ndthis.algid);
566 /* The algorithm here finds all entries for which algorithm is alg and
567 algid matches aldid. We build a list of uuids (keys) linked to that unique
568 algorithm. We then use the id in ndthis as the master*/
569 set<string> matching_ids;
570 matching_ids.insert(ndthis.uuid);
571 /* this approach of pushing iterators to this list that match seemed to
572 be the only way I could make this work correctly. Not sure why, but
573 the added cost over handling this correctly in the loops is small. */
574 std::list<multimap<string, NodeData>::iterator> need_to_erase;
575 for (auto nptr = this->nodes.begin(); nptr != this->nodes.end(); ++nptr) {
576 /* this copy operation is somewhat inefficient, but the cost is small
577 compared to how obscure the code will look if we directly manipulate the
578 second value */
579 NodeData nd(nptr->second);
580 /* this depends upon the distinction between set and multiset. i.e. an
581 insert of a duplicate does nothing*/
582 if ((alg == nd.algorithm) && (algidtest == nd.algid)) {
583 matching_ids.insert(nd.uuid);
584 need_to_erase.push_back(nptr);
585 }
586 }
587 // handle no match situation gracefully
588 if (matching_ids.empty())
589 return string("UNDEFINED");
590 /* Nothing more to do but return the uuid if there is only one*/
591 if (matching_ids.size() == 1)
592 return *(matching_ids.begin());
593 else {
594 for (auto sptr = need_to_erase.begin(); sptr != need_to_erase.end();
595 ++sptr) {
596 nodes.erase(*sptr);
597 }
598 need_to_erase.clear();
599 }
600 /* Here is the complicated case. We use the uuid from ndthis as the master
601 and change all the others. This operation works ONLY because in a multimap
602 erase only invalidates the iterator it points to and others remain valid.
603 */
604 string master_uuid = ndthis.uuid;
605 for (auto sptr = matching_ids.begin(); sptr != matching_ids.end(); ++sptr) {
606 /* Note this test is necessary to stip the master_uuid - no else needed*/
607 if ((*sptr) != master_uuid) {
608 multimap<string, NodeData>::iterator nl, nu;
609 nl = this->nodes.lower_bound(*sptr);
610 nu = this->nodes.upper_bound(*sptr);
611 for (auto nptr = nl; nptr != nu; ++nptr) {
612 NodeData nd;
613 nd = (nptr->second);
614 need_to_erase.push_back(nptr);
615 nodes.insert(pair<string, NodeData>(master_uuid, nd));
616 }
617 }
618 }
619 for (auto sptr = need_to_erase.begin(); sptr != need_to_erase.end(); ++sptr) {
620 nodes.erase(*sptr);
621 }
622
623 return master_uuid;
624}

References mspass::utility::NodeData::algid, mspass::utility::NodeData::algorithm, current_nodedata(), is_empty(), nodes, and mspass::utility::NodeData::uuid.

◆ clear()

void mspass::utility::ProcessingHistory::clear ( )

Clear this history chain - use with caution.

644 {
645 nodes.clear();
646 current_status = ProcessingStatus::UNDEFINED;
647 current_stage = 0;
648 mytype = AtomicType::UNDEFINED;
649 algorithm = "UNDEFINED";
650 algid = "UNDEFINED";
651}

References nodes.

◆ created_by()

std::pair< std::string, std::string > mspass::utility::ProcessingHistory::created_by ( ) const
inline

Return the algorithm name and id that created current node.

605 {
606 std::pair<std::string, std::string> result(algorithm, algid);
607 return result;
608 }

◆ current_nodedata()

NodeData mspass::utility::ProcessingHistory::current_nodedata ( ) const

Return all the attributes of current.

This is a convenience method strictly for the C++ interface (it too nonpythonic to be useful to wrap for python). It returns a NodeData class containing the attributes of the head of the chain. Like the getters above that is needed to save that data.

672 {
673 NodeData nd;
674 nd.status = current_status;
675 nd.uuid = current_id;
676 nd.type = mytype;
677 nd.stage = current_stage;
678 nd.algorithm = algorithm;
679 nd.algid = algid;
680 return nd;
681}

References mspass::utility::NodeData::algid, mspass::utility::NodeData::algorithm, mspass::utility::NodeData::stage, mspass::utility::NodeData::status, mspass::utility::NodeData::type, and mspass::utility::NodeData::uuid.

◆ get_nodes()

multimap< string, NodeData > mspass::utility::ProcessingHistory::get_nodes ( ) const

Retrieve the nodes multimap that defines the tree stucture branches.

This method does more than just get the protected multimap called nodes. It copies the map and then pushes the "current" contents to the map before returning the copy. This allows the data defines as current to not be pushed into the tree until they are needed.

625 {
626 /* Return empty map if it has no data - necessary or the logic
627 below will insert an empty head to the chain. */
628 if (this->is_empty())
629 return nodes; // a way to return an empty container
630 /* This is wrong, I think, but retained to test before removing.
631 remove this once current idea is confirmed. Note if that
632 proves true we can also remove the two lines above as they do
633 nothing useful*/
634 /*
635 NodeData nd;
636 nd=this->current_nodedata();
637 pair<string,NodeData> pn(current_id,nd);
638 multimap<string,NodeData> result(this->nodes);
639 result.insert(pn);
640 return result;
641 */
642 return nodes;
643}

References is_empty(), and nodes.

◆ id()

std::string mspass::utility::ProcessingHistory::id ( ) const
inline

Return the id of this object set for this history chain.

We maintain the uuid for a data object inside this class. This method fetches the string representation of the uuid of this data object.

603{ return current_id; };

◆ inputs()

list< NodeData > mspass::utility::ProcessingHistory::inputs ( const std::string  id_to_find) const

Return a list of data that define the inputs to a give uuids.

This low level getter returns the NodeData objects that define the inputs to the uuid of some piece of data that was used as input at some stage for the current object.

Parameters
id_to_findis the uuid for which input data is desired.
Returns
list of NodeData that define the inputs. Will silently return empty list if the key is not found.
683 {
684 list<NodeData> result;
685 // Return empty list immediately if key not found
686 if (nodes.count(id_to_find) <= 0)
687 return result;
688 /* Note these have to be const_iterators because method is tagged const*/
689 multimap<string, NodeData>::const_iterator upper, lower;
690 lower = nodes.lower_bound(id_to_find);
691 upper = nodes.upper_bound(id_to_find);
692 multimap<string, NodeData>::const_iterator mptr;
693 for (mptr = lower; mptr != upper; ++mptr) {
694 result.push_back(mptr->second);
695 }
696 return result;
697};

References nodes.

◆ is_empty()

bool mspass::utility::ProcessingHistory::is_empty ( ) const

Return true if the processing chain is empty.

This method provides a standard test for an invalid, empty processing chain. Constructors except the copy constructor will all put this object in an invalid state that will cause this method to return true. Only if the chain is initialized properly with a call to set_as_origin will this method return a false.

121 {
122 if ((current_status == ProcessingStatus::UNDEFINED) && (nodes.empty()))
123 return true;
124 return false;
125}

References nodes.

◆ is_origin()

bool mspass::utility::ProcessingHistory::is_origin ( ) const

Return true if the current data is in state defined as "origin" - see class description

132 {
133 if (current_status == ProcessingStatus::RAW ||
134 current_status == ProcessingStatus::ORIGIN)
135 return true;
136 else
137 return false;
138}

◆ is_raw()

bool mspass::utility::ProcessingHistory::is_raw ( ) const

Return true if the current data is in state defined as "raw" - see class description

126 {
127 if (current_status == ProcessingStatus::RAW)
128 return true;
129 else
130 return false;
131}

◆ is_saved()

bool mspass::utility::ProcessingHistory::is_saved ( ) const

Return true if the current data is in state defined as "saved" - see class description

145 {
146 if (current_status == ProcessingStatus::SAVED)
147 return true;
148 else
149 return false;
150}

◆ is_volatile()

bool mspass::utility::ProcessingHistory::is_volatile ( ) const

Return true if the current data is in state defined as "volatile" - see class description

139 {
140 if (current_status == ProcessingStatus::VOLATILE)
141 return true;
142 else
143 return false;
144}

◆ map_as_saved()

string mspass::utility::ProcessingHistory::map_as_saved ( const std::string  alg,
const std::string  algid,
const AtomicType  typ 
)

Prepare the current data for saving.

Saving data is treated as a special form of map operation. That is because a save by our definition is always a one-to-one operation with an index entry for each atomic object. This method pushes a new entry in the history chain tagged by the algorithm/algid field for the writer. It differs from new_map in the important sense that the uuid is not changed. The record this sets in the nodes multimap will then have the same uuid for the key as the that in NodeData. That along with the status set SAVED can be used downstream to recognize save records.

It is VERY IMPORTANT for use of this method to realize this method saves nothing. It only preps the history chain data so calls that follow will retrieve the right information to reconstruct the full history chain. Writers should follow this sequence:

  1. call map_as_saved with the writer name for algorithm definition
  2. save the data and history chain to MongoDB.
  3. be sure you have a copy of the uuid string of the data just saved and call the clear method.
  4. call the set_as_origin method using the uuid saved with the algorithm/id the same as used for earlier call to map_as_saved. This makes the put ProcessingHistory in a state identical to that produced by a reader.
Parameters
algis the algorithm names to assign to the ouput. This would normally be name defining the writer.
algidis an id designator to uniquely define an instance of algorithm. Note that algid must itself be a unique keyword or the history chains will get scrambled. alg is mostly carried as baggage to make output more easily comprehended without additional lookups. Note one model to distinguish records of actual save and redefinition of the data as an origin (see above) is to use a different id for the call to map_as_saved and later call to set_as_origin. This code doesn't care, but that is an implementation detail in how this will work with MongoDB.
typdefines the data type (C++ class) that was just saved.
409 {
410 if (this->is_empty()) {
411 stringstream ss;
412 ss << "Attempt to call this method on an empty history chain for uuid="
413 << this->id() << endl
414 << "Cannot preserve history for writer=" << alg << " with id=" << algid
415 << endl;
416 elog.log_error("ProcessingHistory::map_as_saved", ss.str(),
417 ErrorSeverity::Complaint);
418 return current_id;
419 }
420 /* This is essentially pushing current data to the end of the history chain
421 but using a special id that may or may not be saved by the caller.
422 We use a fixed keyword defined in ProcessingHistory.h assuming saves
423 are always a one-to-one operation (definition of atomic really)*/
424 NodeData nd(this->current_nodedata());
425 pair<string, NodeData> pn(SAVED_ID_KEY, nd);
426 this->nodes.insert(pn);
427 /* Now we reset current to define it as the saver. Then calls to the
428 getters for the multimap will properly insert this data as the end of the
429 chain. Note a key difference from new_map is we don't create a new uuid.
430 I don't think that will cause an ambiguity, but it might be better to
431 just create a new one here - will do it this way unless that proves a problem
432 as the equality of the two might be a useful test for other purposes */
433 algorithm = alg;
434 algid = algid_in;
435 current_status = ProcessingStatus::SAVED;
436 current_id = SAVED_ID_KEY;
437 if (current_stage >= 0)
438 ++current_stage;
439 else {
441 "ProcessingHistory::map_as_saved",
442 "current_stage on entry had not been initialized\nImproper usage will "
443 "create an invalid history chain that may cause downstream problems",
444 ErrorSeverity::Complaint);
445 current_stage = 0;
446 }
447 mytype = typ;
448 return current_id;
449}
std::string id() const
Definition ProcessingHistory.h:603

References current_nodedata(), elog, id(), is_empty(), mspass::utility::ErrorLogger::log_error(), and nodes.

◆ merge()

void mspass::utility::ProcessingHistory::merge ( const ProcessingHistory &  data_to_add)

Merge the history nodes from another.

Parameters
data_to_addis the ProcessingHistory of the data object to be merged.
452 {
453
454 if (data_to_add.is_empty()) {
455 stringstream ss;
456 ss << "Data with uuid=" << data_to_add.id() << " has an empty history chain"
457 << endl
458 << "At best this will leave ProcessingHistory incomplete" << endl;
459 elog.log_error("ProcessingHistory::merge", ss.str(),
460 ErrorSeverity::Complaint);
461 } else {
462 multimap<string, NodeData>::iterator nptr;
463 multimap<string, NodeData> newhistory = data_to_add.get_nodes();
464 multimap<string, NodeData>::iterator nl, nu;
465 for (nptr = newhistory.begin(); nptr != newhistory.end(); ++nptr) {
466 string key(nptr->first);
467 /* if the data_to_add's key matches its current id,
468 we merge all the nodes under the current id of *this. */
469 if (key == data_to_add.current_id) {
470 this->nodes.insert(std::make_pair(this->current_id, nptr->second));
471 } else if (this->nodes.count(key) > 0) {
472 nl = this->nodes.lower_bound(key);
473 nu = this->nodes.upper_bound(key);
474 for (auto ptr = nl; ptr != nu; ++ptr) {
475 NodeData ndtest(ptr->second);
476 if (ndtest != (nptr->second)) {
477 this->nodes.insert(*nptr);
478 }
479 }
480 } else {
481 this->nodes.insert(*nptr);
482 }
483 }
484 }
485}

References elog, get_nodes(), id(), is_empty(), mspass::utility::ErrorLogger::log_error(), and nodes.

◆ new_ensemble_process()

string mspass::utility::ProcessingHistory::new_ensemble_process ( const std::string  alg,
const std::string  algid,
const AtomicType  typ,
const std::vector< ProcessingHistory * >  parents,
const bool  create_newid = true 
)

Define history chain for an algorithm with multiple inputs in an ensemble.

Use this method to define the history chain for an algorithm that has multiple inputs for each output. Each output needs to call this method to build the connections that define how all inputs link to the the new data being created by the algorithm that calls this method. Use this method for map operators that have an ensemble object as input and a single data object as output. This method should be called in creation of the output object. If the algorthm builds multiple outputs to build an output ensemble call this method for each output before pushing it to the output ensemble container.

This method should not be used for a reduce operation in spark. It does not satisfy the associative rule for reduce. Use accumulate for reduce operations.

Normally, it makes sense to have the boolean create_newid true so it is guaranteed the current_id is unique. There is little cost in creating a new one if there is any doubt the current_id is not a duplicate. The false option is there only for rare cases where the current id value needs to be preserved.

Note the vector of data passed is raw pointers for efficiency to avoid excessive copying. For normal use this should not create memory leaks but make sure you don't try to free what the pointers point to or problems are guaranteed. It is VERY IMPORTANT to realize that all the pointers are presumed to point to the ProcessingHistory component of a set of larger data object (Seismogram or TimeSeries). The parents do not all have be a common type as if they have valid history data within them their current type will be defined.

This method ALWAYS marks the status as VOLATILE.

Parameters
algis the algorithm names to assign to the origin node. This would normally be name defining the algorithm that makes sense to a human.
algidis an id designator to uniquely define an instance of algorithm. Note that algid must itself be a unique keyword or the history chains will get scrambled. alg is mostly carried as baggage to make output more easily comprehended without additional lookups.
typdefines the data type (C++ class) the algorithm that is generating this data will create.
parentsis a vector of ProcessingHistory pointers for all input data objects used to create this ensemble.
create_newidis a boolean defining how the current id is handled. As described above, if true the method will call newid and set that as the current id of this data object. If false the current value is left intact.
Returns
a string representation of the uuid of the data to which this ProcessingHistory is now attached.
186 {
187 if (create_newid) {
188 this->newid();
189 }
190 /* We need to clear the tree contents because all the parents will
191 branch from this. Hence, we have to put the node data into an empty
192 container */
193 this->clear();
194 algorithm = alg;
195 algid = algid_in;
196 mytype = typ;
197 /* Initialize current stage but assume it will be updated as max of
198 parents below */
199 current_stage = 0;
200 multimap<string, NodeData>::const_iterator nptr, nl, nu;
201 size_t i;
202 /* current_stage can be ambiguous from multiple inputs. We define
203 the current stage from a reduce as the largest stage value found
204 in all inputs. Note we only test the stage value at the head for
205 each parent */
206 int max_stage(0);
207 for (i = 0; i < parents.size(); ++i) {
208 if (parents[i]->is_empty()) {
209 stringstream ss;
210 ss << "Vector member number " << i << " with uuid=" << parents[i]->id()
211 << " has an empty history chain" << endl
212 << "At best the processing history data will be incomplete" << endl;
213 elog.log_error("ProcessingHistory::new_ensemble_process", ss.str(),
214 ErrorSeverity::Complaint);
215 continue;
216 }
217 multimap<string, NodeData> parent_node_data(parents[i]->get_nodes());
218 /* We also have to get the head data with this method now */
219 NodeData nd = parents[i]->current_nodedata();
220 if (nd.stage > max_stage)
221 max_stage = nd.stage;
222 for (nptr = parent_node_data.begin(); nptr != parent_node_data.end();
223 ++nptr) {
224 /*Adding to nodes multimap has a complication. It is possible in
225 some situations to have duplicate node data coming from different
226 inputs. The method we use to reconstruct the processing history tree
227 will be confused by such duplicates so we need to test for pure
228 duplicates in NodeData values. This algorithm would not scale well
229 if the number of values with a common key is large for either
230 this or parent[i]*/
231 string key(nptr->first);
232 if (this->nodes.count(key) > 0) {
233 nl = this->nodes.lower_bound(key);
234 nu = this->nodes.upper_bound(key);
235 for (auto ptr = nl; ptr != nu; ++ptr) {
236 NodeData ndtest(ptr->second);
237 if (ndtest != (nptr->second)) {
238 this->nodes.insert(*nptr);
239 }
240 }
241 } else {
242 /* No problem just inserting a node if there were no previous
243 entries*/
244 this->nodes.insert(*nptr);
245 }
246 }
247 /* Also insert the head data */
248 pair<string, NodeData> pnd(current_id, nd);
249 this->nodes.insert(pnd);
250 }
251 current_stage = max_stage;
252 /* Now reset the current contents to make it the base of the history tree.
253 Be careful of uninitialized current_stage*/
254 if (current_stage >= 0)
255 ++current_stage;
256 else {
257 elog.log_error("ProcessingHistory::new_ensemble_process",
258 "current_stage for none of the parents was "
259 "initialized\nImproper usage will create an invalid history "
260 "chain that may cause downstream problems",
261 ErrorSeverity::Complaint);
262 current_stage = 0;
263 }
264 algorithm = alg;
265 algid = algid_in;
266 // note this is output type - inputs can be variable and defined by nodes
267 mytype = typ;
268 current_status = ProcessingStatus::VOLATILE;
269 return current_id;
270}
void clear()
Definition ProcessingHistory.cc:644
std::multimap< std::string, mspass::utility::NodeData > get_nodes() const
Definition ProcessingHistory.cc:625

References clear(), elog, get_nodes(), is_empty(), mspass::utility::ErrorLogger::log_error(), newid(), nodes, and mspass::utility::NodeData::stage.

◆ new_map() [1/2]

std::string mspass::utility::ProcessingHistory::new_map ( const std::string  alg,
const std::string  algid,
const AtomicType  typ,
const ProcessingHistory &  data_to_clone,
const ProcessingStatus  newstatus = ProcessingStatus::VOLATILE 
)

Define this algorithm as a one-to-one map.

Many algorithms define a one-to-one map where each one input data object creates one output data object. This class allows the input and output to be different data types requiring only that one input will map to one output. It differs from the overloaded method with fewer arguments in that it should be used if you need to clear and refresh the history chain for any reason. Known examples are creating simulation waveforms for testing within a workflow that have no prior history data loaded but which clone some properties of another piece of data. This method should be used in any situation where the history chain in the current data is wrong but the contents are the linked to some other process chain. It is supplied to cover odd cases, but use will likely be rare.

Parameters
algis the algorithm names to assign to the origin node. This would normally be name defining the algorithm that makes sense to a human.
algidis an id designator to uniquely define an instance of algorithm. Note that algid must itself be a unique keyword or the history chains will get scrambled. alg is mostly carried as baggage to make output more easily comprehended without additional lookups.
typdefines the data type (C++ class) the algorithm that is generating this data will create.
data_to_cloneis reference to the ProcessingHistory section of a parent data object that should be used to override the existing history chain.
newstatusis how the status marking for the output. Normal (default) would be VOLATILE. This argument was included mainly for flexibility in case we wanted to extend the allowed entries in ProcessingStatus.
374 {
375 /* We must be sure the chain is empty before we push the clone's data there*/
376 this->clear();
377 /* this works because get_nodes pushes the current data to the nodes
378 multimap. We intentionally do not test for an empty nodes map
379 assuming one wouldn't call this without knowing that was necessary.
380 That may be an incorrect assumption, but will use it until proven otherwise*/
381 nodes = copy_to_clone.get_nodes();
382 NodeData nd;
383 nd = this->current_nodedata();
384 /* We always need a new id here for this object we are handling as the child
385 */
386 current_id = this->newid();
387 pair<string, NodeData> pn(current_id, nd);
388 this->nodes.insert(pn);
389 algorithm = alg;
390 algid = algid_in;
391 current_status =
392 newstatus; // Probably should default in include file to VOLATILE
393 if (current_stage >= 0)
394 ++current_stage;
395 else {
397 "ProcessingHistory::new_map",
398 "current_stage on entry had not been initialized\nImproper usage will "
399 "create an invalid history chain that may cause downstream problems",
400 ErrorSeverity::Complaint);
401 current_stage = 0;
402 }
403 mytype = typ;
404 return current_id;
405}

References clear(), current_nodedata(), elog, get_nodes(), mspass::utility::ErrorLogger::log_error(), newid(), and nodes.

◆ new_map() [2/2]

std::string mspass::utility::ProcessingHistory::new_map ( const std::string  alg,
const std::string  algid,
const AtomicType  typ,
const ProcessingStatus  newstatus = ProcessingStatus::VOLATILE 
)

Define this algorithm as a one-to-one map of same type data.

Many algorithms define a one-to-one map where each one input data object creates one output data object. This (overloaded) version of this method is most appropriate when input and output are the same type and the history chain (ProcessingHistory) is what the new algorithm will alter to make the result when it finishes. Use the overloaded version with a separate ProcessingHistory copy if the current object's data are not correct. In this algorithm the chain for this algorithm is simply appended with new definitions.

Parameters
algis the algorithm names to assign to the origin node. This would normally be name defining the algorithm that makes sense to a human.
algidis an id designator to uniquely define an instance of algorithm. Note that algid must itself be a unique keyword or the history chains will get scrambled. alg is mostly carried as baggage to make output more easily comprehended without additional lookups.
typdefines the data type (C++ class) the algorithm that is generating this data will create.
newstatusis how the status marking for the output. Normal (default) would be VOLATILE. This argument was included mainly for flexibility in case we wanted to extend the allowed entries in ProcessingStatus.
333 {
334 if (this->is_empty()) {
335 stringstream ss;
336 ss << "Attempt to call this method on an empty history chain for uuid="
337 << this->id() << endl
338 << "Cannot preserve history for algorithm=" << alg
339 << " with id=" << algid << endl;
340 elog.log_error("ProcessingHistory::new_map", ss.str(),
341 ErrorSeverity::Complaint);
342 return current_id;
343 }
344 /* In this case we have to push current data to the history chain */
345 NodeData nd;
346 nd = this->current_nodedata();
347 /* We always need a new id here for this object we are handling as the child
348 */
349 current_id = this->newid();
350 /* The new id is now the key to link back to previous record so we insert
351 nd with the new key to define that link */
352 pair<string, NodeData> pn(current_id, nd);
353 this->nodes.insert(pn);
354 algorithm = alg;
355 algid = algid_in;
356 current_status =
357 newstatus; // Probably should default in include file to VOLATILE
358 if (current_stage >= 0)
359 ++current_stage;
360 else {
362 "ProcessingHistory::new_map",
363 "current_stage on entry had not been initialized\nImproper usage will "
364 "create an invalid history chain that may cause downstream problems",
365 ErrorSeverity::Complaint);
366 current_stage = 0;
367 }
368 mytype = typ;
369 return current_id;
370}

References current_nodedata(), elog, id(), is_empty(), mspass::utility::ErrorLogger::log_error(), newid(), and nodes.

◆ newid()

string mspass::utility::ProcessingHistory::newid ( )

Create a new id.

This creates a new uuid - how is an implementation detail but here we use boost's random number generator uuid generator that has some absurdly small probability of generating two equal ids. It returns the string representation of the id created.

664 {
665 boost::uuids::random_generator gen;
666 boost::uuids::uuid uuidval;
667 uuidval = gen();
668 this->current_id = boost::uuids::to_string(uuidval);
669 return current_id;
670}

◆ number_inputs() [1/2]

int mspass::utility::ProcessingHistory::number_inputs ( ) const

Return the number of inputs used to create current data.

In a number of contexts it can be useful to know the number of inputs defined for the current object. This returns that count.

661 {
662 return this->number_inputs(current_id);
663}
int number_inputs() const
Definition ProcessingHistory.cc:661

References number_inputs().

◆ number_inputs() [2/2]

int mspass::utility::ProcessingHistory::number_inputs ( const std::string  uuidstr) const

Return the number of inputs defined for any data in the process chain.

This overloaded version of number_inputs asks for the number of inputs defined for an arbitrary uuid. This is useful only if backtracing the ancestory of a child.

Parameters
uuidstris the uuid string to check in the ancestory record.
655 {
656 // Return result is int to mesh better with python even though
657 // count returns size_t
658 int n = nodes.count(testuuid);
659 return n;
660}

References nodes.

◆ number_of_stages()

size_t mspass::utility::ProcessingHistory::number_of_stages ( )
overridevirtual

Return number of processing stages that have been applied to this object.

One might want to know how many processing steps have been previously applied to produce the current data. For linear algorithms that would be useful only in debugging, but for an iterative algorithm it can be essential to avoid infinite loops with a loop limit parameter. This method returns how many times something has been done to alter the associated data. It returns 0 if the data are raw.

Important note is that the number return is the number of processing steps since the last save. Because a save operation is assumed to save the history chain then flush it there is not easy way at present to keep track of the total number of stages. If we really need this functionality it could be easily retrofitted with another private variable that is not reset when the clear method is called.

Reimplemented from mspass::utility::BasicProcessingHistory.

151{ return current_stage; }

◆ operator=()

ProcessingHistory & mspass::utility::ProcessingHistory::operator= ( const ProcessingHistory &  parent)

Assignment operator.

700 {
701 if (this != (&parent)) {
703 nodes = parent.nodes;
704 current_status = parent.current_status;
705 current_id = parent.current_id;
706 current_stage = parent.current_stage;
707 mytype = parent.mytype;
708 algorithm = parent.algorithm;
709 algid = parent.algid;
710 elog = parent.elog;
711 }
712 return *this;
713}
BasicProcessingHistory & operator=(const BasicProcessingHistory &parent)
Definition ProcessingHistory.h:95

References elog, nodes, and mspass::utility::BasicProcessingHistory::operator=().

◆ set_as_origin()

void mspass::utility::ProcessingHistory::set_as_origin ( const std::string  alg,
const std::string  algid,
const std::string  uuid,
const AtomicType  typ,
bool  define_as_raw = false 
)

Set to define this as the top origin of a history chain.

This method should be called when a new object is created to initialize the history as an origin. Note again an origin may be raw but not all origins are define as raw. This interface controls that through the boolean define_as_raw (false by default). python wrappers should define an alternate set_as_raw method that calls this method with define_as_raw set true.

It is VERY IMPORTANT to realize that the uuid argument passed to this method is if fundamental importance. That string is assumed to be a uuid that can be linked to either a parent data object read from storage and/or linked to a history chain saved by a prior run. It becomes the current_id for the data to which this object is a parent. This method also always does two things that define how the contents can be used. current_stage is ALWAYS set 0. We distinguish a pure origin from an intermediate save ONLY by the status value saved in the history chain. That is, only uuids with status set to RAW are viewed as guaranteed to be stored. A record marked ORIGIN is assumed to passed through save operation. To retrieve the history chain from multiple runs the pieces have to be pieced together by history data stored in MongoDB.

The contents of the history data structures should be empty when this method is called. That would be the norm for any constructor except those that make a deep copy. If unsure the clear method should be called before this method is called. If it isn't empty it will be cleared anyway and a complaint message will be posted to elog.

Parameters
algis the algorithm names to assign to the origin node. This would normally be a reader name, but it could be a synthetic generator.
algidis an id designator to uniquely define an instance of algorithm. Note that algid must itself be a unique keyword or the history chains will get scrambled.
uuidunique if for this data object (see note above)
typdefines the data type (C++ class) "this" points to. It might be possible to determine this dynamically, but a design choice was to only allow registered classes through this mechanism. i.e. the enum class typ implements has a finite number of C++ classes it accepts. The type must be a child ProcessingHistory.
define_as_rawsets status as RAW if true and ORIGIN otherwise.
Exceptions
Neverthrows an exception BUT this method will post a complaint to elog if the history data structures are not empty and it the clear method needs to be called internally.
163 {
164 const string base_error("ProcessingHistory::set_as_origin: ");
165 if (nodes.size() > 0) {
166 elog.log_error(alg + ":" + algid_in,
167 base_error + "Illegal usage. History chain was not empty. "
168 " Calling clear method and continuing",
169 ErrorSeverity::Complaint);
170 this->clear();
171 }
172 if (define_as_raw) {
173 current_status = ProcessingStatus::RAW;
174 } else {
175 current_status = ProcessingStatus::ORIGIN;
176 }
177 algorithm = alg;
178 algid = algid_in;
179 current_id = uuid;
180 mytype = typ;
181 /* Origin/raw are always defined as stage 0 even after a save. */
182 current_stage = 0;
183}

References clear(), elog, mspass::utility::ErrorLogger::log_error(), and nodes.

◆ set_id()

void mspass::utility::ProcessingHistory::set_id ( const std::string  newid)

Set the uuid manually.

It may occasionally be necessary to create a uuid by some other mechanism. This allows that, but this method should be used with caution and only if you understand the consequences.

Parameters
newidis string definition to use for the id.
671{ this->current_id = newid; }

References newid().

◆ stage()

int mspass::utility::ProcessingHistory::stage ( ) const
inline

Return the current stage count for this object.

We maintain a counter of the number of processing steps that have been applied to produce this data object. This simple method returns that counter. With this implementation this is identical to number_of_stages. We retain it in the API in the event we want to implement an accumulating counter.

595{ return current_stage; };

◆ status()

ProcessingStatus mspass::utility::ProcessingHistory::status ( ) const
inline

Return the current status definition (an enum).

597{ return current_status; };

Member Data Documentation

◆ elog

ErrorLogger mspass::utility::ProcessingHistory::elog

Error log for non-fatal processing-history consistency complaints.

◆ nodes

std::multimap<std::string, mspass::utility::NodeData> mspass::utility::ProcessingHistory::nodes
protected

Connections between each data object uuid and its input node records.

The key is the uuid of a data object, and each associated NodeData value describes one input used to create the data identified by that uuid.


The documentation for this class was generated from the following files: